A search engine returns documents ranked by relevance and leaves you to work out what is true. A language model returns fluent prose with no accountable relationship between what it asserts and what any document says. Between finding and generating sits a third thing, and it is the one knowledge work actually needs.
Epago is a decentralized network in which thousands of competing AI researchers discover, verify and synthesize knowledge — every claim traced to a source, every improvement proven, every result replayable by a stranger from public data, forever. The output is not a document to read and not a guess to trust. It is verified intelligence, and the purpose of the network is to make it open, cheap and continuous.
We are not building one AI researcher. We are building the market where thousands compete, and the winner has to prove it.
They will be other AI systems. That audience changes what an answer has to be. A person can read a paragraph and decide how much to believe it; an agent cannot. An agent cannot act on an answer it cannot verify, and it cannot pay for one it cannot check.
So machine-readable output, source-traceable claims, and a verification trail the caller can re-derive without trusting the responder are not conveniences layered on at the end. They are the requirement — and they are why this network puts verification at the base of its stack rather than above it.
This is where the numbers live. A bigger model is mostly a bigger memory of the world — but a deep-research agent is not asked to remember, it is asked to find out: break a question into sub-questions, choose sources, read them, notice that two studies disagree, and attach every number it reports to the document it came from. Procedures distill into small models. Memorized world-knowledge does not. Once the corpus supplies the facts, the deciding variable is no longer size.
To be exact about the claim: this is superiority at deep research, per unit of compute. Nothing here asserts that a ~3.3B-active model out-thinks a frontier model in the open field, and Epago never will.
These are the published results of the genesis base model — Tongyi-DeepResearch-30B-A3B, a 30.5B mixture-of-experts that activates about 3.3B parameters per token, with a 128K-token context. It outperforms OpenAI o3 and DeepSeek-V3.1 (671B) on 5 of the 7 agentic deep-research benchmarks reported. o3 still leads both BrowseComp variants — 49.7 vs 43.4 in English, 58.1 vs 46.7 in Chinese — and both columns are printed below rather than dropped, because selective reporting is precisely the failure this protocol exists to remove. Epago does not inherit these scores as its own: every crown since has had to beat them.
| MODEL | SIZE | {{ b.name }} |
|---|---|---|
| {{ r.name }} | {{ r.size }} | {{ c.v }} |
Deep research is the engine, not the product. What a research network ultimately supplies is not answers but knowledge that has been checked and stays checked. The architecture for that is a stack, six layers deep, each one inheriting the mechanical grading and replayable verdicts of the layer beneath it.
This is the one place on this page where live and planned are labelled explicitly, and it is the only time it needs saying: nothing marked THE HORIZON is live, earns emissions, or is assumed in any number anywhere on this page. Everything marked RUNNING TODAY exists as running software. That separation is what makes the rest of the stack credible rather than promotional.
The mechanism is domain-general by construction, because grading never asks what is true — only whether an answer matches a key already proven to exist in a pinned corpus. That makes it usable in the scientific literature, in law, in patents, in regulatory filings, in financial disclosure. What those fields share is that rigorous, source-traceable work is slow and expensive today, and that the buyers who need it most are the ones privacy rules and high-risk-AI regulation lock out of hosted AI. These are illustrative candidates, not a committed schedule and not an ordering: sequencing is community-governed, and generations run in parallel.
A domain is a chain contract: a corpus, a set of task templates, an architecture pin. Adding one does not fork the engine — it starts another generation, with its own king, its own miners and its own crown, running in parallel on the same machinery. Two invariants decide what is admissible: answers must be checkable against the corpus without an opinion, and fresh documents must arrive on a schedule. Four axes compound at once.
Duels run against a pinned scientific-paper corpus, searched locally with BM25. No live internet, no API fees, byte-identical on every validator's machine. That is what makes the evaluation reproducible. And the corpus is part of the chain contract rather than the code: a second generation aimed at case law or patents changes the configuration, not the engine.
Every task is built on one rule: the question never contains a string that gives the source away, and the answer can never be copied out of a search result — it has to be reached by real research, then checked mechanically against a key the validator re-derives from the corpus. The exam comes in three tiers, one shape.
A challenger is tested head-to-head against the reigning king: both answer the identical ~1,000 tasks, and only the per-task difference is scored, so exam difficulty cancels out. 800 are drawn from a pool sealed and committed by digest in advance, selected by a block hash that did not exist until the challenger was frozen — and a question once published is retired, never asked again; 200 more come from each validator's private holdout of freshly published papers.
A lucky exam can't crown anyone. Acceptance turns on a statistical test, not a raw score — and the crown only changes hands when independent validators agree. No company decides; coronation is a pure function of what is on the chain.
A duel does not only produce a verdict. Every task inside it emits a verified record — the task, the model's answer, and whether that answer was right — over documents published after the models were trained. That is verified reward data: the scarcest input in post-training, and the competition generates it as a side effect of doing its job.
It cannot be scraped, because it does not exist anywhere else. There is no public dump of graded research episodes over documents that postdate the model reading them. A competitor who wanted this corpus would have to run the machine for a year to accumulate it — and by then the machine has another year of it.
And the exam itself compounds the same way: when a round closes, its private questions are published. Week after week that deposits a public deep-research benchmark that did not exist before — fresh, verified, never reused — with Epago's name on every release.
These four steps are the whole procedure — and not one of them requires the answer to be sitting in the weights. That is the reason a model this small can do this work.
The open deep-research frontier is already small and already strong. Its problem is that it stops moving the day a lab ships the checkpoint and turns to the next thing — and that when a successor does arrive, nothing tells you whether it is genuinely better or merely tuned to the benchmark it is announced with.
Epago replaces that with a competition. Anyone in the world can train a better version; independent validators test each challenger against the reigning model on fresh questions chosen after the challenger's weights were frozen. Only a statistically proven improvement — 99.9% confidence — takes the crown and the rewards.
No company decides. Nothing can be faked. Every result is publicly replayable. That is the moat: while a closed lab ships when it chooses to, this model improves whenever anyone on Earth finds a way to make it better — and every round of it deposits reward data nobody else has.
This is why it runs on a chain rather than on a company's leaderboard. Every anti-gaming property collapses the moment the evaluator and the beneficiary are the same party — which is precisely how one provider came to test 27 model variants privately and publish only the best. Emissions buy measured capability, not claims: no participant, the Epago Foundation included, can quietly become the judge.