Tessera is the curation engine for training sets — dedupe, score, and compose corpora that make the same model 4–9 points better on the same compute.
Web-scale corpora are 40% near-duplicates and noise. Compute is priced to the penny while the thing it trains on goes unmeasured.
benchmark improvement from curation alone at fixed compute, replicated across three model families in our published evals (2025) — cheaper than any scaling law says compute can buy.
Two frontier labs run Tessera inside their own VPCs — the deals that anchor the enterprise tier.
Quality models distilled from 400+ customer eval suites — accuracy compounds cross-customer.
100B-doc dedupe in 9 hours; rivals quote weeks.
Our per-token provenance format is entering two compliance frameworks.
“Same cluster, same architecture, Tessera mixture: +6.2 on our eval suite. That is a result we had budgeted eight figures of compute to reach.”
Compute got its platforms. Serving is getting its runtimes. Data — the highest-leverage layer — is still up for grabs.