One engine, three stages, no dead ends
Traditional selection treats design, synthesis and measurement as separate projects handed between teams. We run them as a single instrumented loop in which every measurement updates the model that produced the candidate.
Agentic hypothesis & target modeling
A scientific agent acts as the reasoning orchestrator for a campaign. Given a target protein or structured RNA, it assembles what is already known before proposing anything new: published binders and their modalities, structural data and its resolution, known allosteric or regulatory sites, counter-targets in the same family, and the assay formats others have used.
From that the agent proposes candidate binding surfaces and, critically, writes down the design constraints as an explicit, reviewable specification. Constraints are not preferences. They are the objective the downstream models optimize against.
A scientist signs off on that specification before generation starts. The agent records its reasoning and the evidence behind each constraint, so a disagreement can be traced to a source rather than argued from intuition.
Constraints the agent sets
- Secondary structure motifs: stem length, loop size, required or forbidden folds
- Thermodynamic stability, minimum free energy and ensemble defect targets
- GC content window and homopolymer limits for synthesis quality
- Nuclease resistance strategy and modification budget
- Counter-selection targets and specificity margin
- Immunostimulatory motifs to avoid
Hybrid machine learning design loop
The agent does not generate sequences itself. It calls specialized models through a tool interface and is responsible for deciding what to call, how to interpret disagreement between tools, and which candidates survive.
Generative sequence models propose libraries conditioned on the constraint specification. Structure tools fold each candidate and report minimum free energy, the suboptimal ensemble and how reliably the intended motif forms. Docking and scoring estimate the fit against the proposed binding surface. Candidates that pass structural filters are then scored for developability and synthesis feasibility.
The output is not a single answer. It is a ranked, diverse shortlist with an explicit reason for each selection and a set of deliberately included controls, so that a cycle which fails still produces interpretable information.
Tools in the loop
- RNA language models
- Protein sequence models
- Diffusion & transformer folding
- ViennaRNA / RNAfold
- Secondary structure ensembles
- Molecular docking
- Developability scoring
Closed-loop wet-lab validation
Design is only credible if it is measured. We work with wet labs at KAIST to synthesize shortlisted candidates and run high-throughput binding characterization: surface plasmon resonance and biolayer interferometry for kinetics, and deep sequencing of selection pools where a hybrid selection arm is used.
Readouts return to the agent as structured data, not as a slide deck. It performs post-assay error analysis: separating candidates that failed to bind from those that failed to fold, identifying assay artifacts such as mass transport limitation or nonspecific surface binding, and checking whether misses cluster on a shared motif.
That analysis updates the structure–activity hypotheses and rewrites the constraint specification for the next cycle. The loop closes, and the next library is generated against a better objective.
What comes back
- Association and dissociation rates, not just an endpoint affinity
- Specificity against counter-targets in the same buffer
- Folding and refolding behaviour across magnesium and salt conditions
- Stability in nuclease-containing matrix
- Sequence-level enrichment data where a selection arm is run
What we insist on
A platform that cannot be audited is not a platform. It is a black box with a marketing page.
Every decision carries a reason
Each shortlisted candidate records the constraints it satisfies, the tools that scored it and the hypothesis it tests. Reviewers can disagree with a specific step.
Scientists hold the gates
Agents propose, prioritize and explain. A human approves the constraint specification and the synthesis list. Nothing reaches the bench unreviewed.
Negative results are kept
Non-binders are labeled data. They are stored with full assay context and reused, rather than discarded at the end of a campaign.
Developability is an objective, not a filter
Stability and manufacturability enter at generation time, so we are not forced to modify a good binder until it stops binding.
Against a conventional campaign
Both approaches end with a characterized binder. They differ in how much is learned on the way there.
| Dimension | Conventional SELEX | Sanghyun engine |
|---|---|---|
| Search strategy | Iterative enrichment of whatever the starting pool happened to contain | Model-guided generation across regions the pool never sampled |
| Sequence space reached | Limited by library synthesis and amplification survival | Limited by scoring confidence, which improves each cycle |
| Developability | Assessed after a lead is chosen | Encoded as constraints before generation |
| Failure handling | Non-binders discarded between rounds | Retained as labeled negatives with assay context |
| Specificity | Addressed by counter-selection rounds | Counter-targets scored in silico, then confirmed on instrument |
| Auditability | Round-level enrichment records | Per-candidate reasoning, constraints and tool scores |
See it run on your target
We scope a first campaign around one target and one clearly defined success criterion, so the result is interpretable either way.