Molecular Discovery Engine & Sealed Catalogue

33.3 million molecular binding predictions. Already generated, cryptographically sealed, available to license today.

Origin Neural built a complete computational discovery catalogue and the engine that produced it. The work is finished, enriched with ADMet and drug-likeness profiling, and signed and anchored to a public blockchain — so you can confirm independently, without our involvement, that this data predates our first conversation and has not been altered since.

The engine is compact enough to run on infrastructure you already own. No supercomputer purchase required.

Sealed 22 February 2026 2,158 batches anchored on-chain
The Asset

What exists today

These are final figures. The catalogue was sealed on 22 February 2026 and has not changed since.

33,325,916Molecular binding predictions
17,195,382Enriched with ADMet profiling
5,929,631Unique molecular scaffolds
2,158Batches anchored on-chain
Why the catalogue is sealed, stated plainly

Origin Neural is bootstrapped. We funded this work ourselves and ran the pipeline until we had a catalogue worth licensing, then stopped generating new predictions because generating at volume is expensive and we chose not to burn capital producing inventory before securing a partner. Model development never stopped — training runs continue and are signed and anchored daily, which anyone can verify on-chain. The engine is intact and restarts on demand, against your targets, on your schedule. Restarting it is one of the things a partnership pays for.

The Engine

Built for efficiency, not for scale spending

The industry trend is toward very large in-house compute. Our approach went the other way: a deliberately small network, trained for geometric stability, that produced this entire catalogue on a bootstrapped budget.

Model size

73,352 parameters

Small enough to run on commodity hardware and deploy inside an existing research environment without new infrastructure or procurement.

Reported accuracy

85.40% holdout

Measured on data held out from training. Self-reported — we recommend independent benchmarking against your internal reference sets before any commitment.

Stability

K3 spectral 0.000016

A near-zero measure of weight-manifold stability across five training cycles, with isometry loss of 0.386 relating to chirality and rotational robustness.

What we are not claiming

A mathematically stable model is not automatically a clinically accurate one, and a computational prediction is not a drug. These metrics describe how the model behaved, not whether a compound will work in an organism. We expect and welcome independent validation.

Commercial Options

What we license

Five distinct offerings. Most partners start with one and expand.

01

Catalogue access

The existing predictions, in full or scoped to a target family, disease area, or scaffold class. Includes compound dossiers, SMILES, ADMet profiles, and scaffold analysis with bulk export.

02

Engine licensing

Deploy the prediction engine inside your own environment and run it against your proprietary targets. Your data never leaves your infrastructure.

03

Commissioned campaigns

Provide a biological target; we restart the pipeline and generate a new, dedicated prediction set to your specification, anchored and delivered.

04

Verification layer

Anchor your own computational results to establish provable priority dates and tamper-evident audit trails — for IP disputes, regulatory records, or multi-party collaborations.

05

Research collaboration

Joint work on a disease area, target class, or compound family, with terms structured around shared discovery rather than a flat data purchase.

Next step

A scoped evaluation

We can prepare a sample set against one target of your choosing so your team can assess quality directly, before commercial terms are discussed.

Full Disclosure

What the catalogue actually contains

A screening funnel only has value if you can see where the filters land. These are the real distributions, published so your team can judge fit before a conversation rather than after one.

33,325,916Total predictions
17,195,382ADMet-enriched
12,150,319Drug-like
295,756VAULT-tier shortlist

Composition and filter pass rates

MeasureValueNote
Peptide-like21,151,570The majority of the catalogue
Drug-like12,150,319Conventional small-molecule space
Fragment24,027Fragment-based starting points
Lipinski compliant31.8%Of enriched compounds
PAINS clean39.5%Free of common assay-interference motifs
Veber pass3.6%Stricter oral-bioavailability criteria
Lead-like0.3%Strictest filter; approximately 100,000 compounds
Avg. synthetic accessibility4.34Scale of 1 (easy) to 10 (hard)
Read this honestly

Roughly two thirds of the catalogue is peptide-like rather than conventional small-molecule, and the strictest lead-like filter passes 0.3%. If your program is exclusively small-molecule oral, the relevant subset is the 12.1 million drug-like compounds, not the headline 33 million. We would rather you know that now than discover it in diligence.

Structural quality of enriched compounds

ClassificationCompoundsMeaning
Green3,108,279Few structural concerns identified
Yellow3,595,212Caution flags identified
Red10,491,891Significant structural or toxicity concerns
Independent Verification

You do not have to take our word for any of this

Every batch was hashed into a Merkle tree, wrapped in an ECDSA-signed payload, and written to the Bitcoin SV blockchain as it was produced. Those records are public and permanent. Anyone can read them without our cooperation.

Timestamped

The block proves the date

Block timestamps are established by network consensus. A record in block 937,482 cannot be backdated afterwards by us or anyone else.

Tamper-evident

The Merkle root proves the contents

Changing any single prediction changes its batch's Merkle root, which then no longer matches the immutable on-chain value. Silent revision is impossible.

Authored

The signature proves the source

Each payload carries an ECDSA signature, so the record was not merely posted to the chain but authored by a specific key holder — and that is checkable.

Verify it in about fifteen seconds

We publish the raw anchors, the exact payload structure, and a dependency-free script that recovers the signing key and checks it against the address the payload claims. Read the script before you run it — it is short enough to audit.

Open the Verification Page
Direct Answers

Questions we expect

Why would we license predictions instead of generating our own?

If you already run large in-house compute, you may not need the catalogue — but you may still want the engine for its efficiency, or the verification layer, which is not something compute alone provides. If you do not run that infrastructure, this catalogue represents work already paid for and independently timestamped, available immediately rather than after a procurement and build cycle.

Why did you stop generating predictions?

Cost. We are bootstrapped and self-funded this work. Rather than continue spending on compute to grow inventory we had not yet monetised, we sealed the catalogue and turned to partnerships. The engine is intact and restarts on demand.

Has the engine been independently benchmarked?

Not yet by a third party. Our reported holdout accuracy is 85.40%. We consider independent benchmarking against a partner's internal reference sets a reasonable precondition for any significant agreement, and we will support and fund it.

What does a binding score actually mean?

It is a model-based estimate of interaction strength with a biological target, intended for ranking and prioritisation. It is not evidence of medical effectiveness and should not be read as such.

Can we evaluate quality before committing?

Yes. We will prepare a scoped sample against a target you nominate so your computational chemists can assess it directly. We would rather be evaluated on output than on claims.

Who owns what the engine produces for us?

Negotiable and defined per agreement. For commissioned campaigns and engine deployments run against your proprietary targets, our expectation is that the outputs are yours.

Next Step

Start with a scoped evaluation

Tell us a target and what your program needs. We will come back with a sample set and an honest assessment of whether this catalogue fits — including if it does not.