Tessera
Series A · 2026
Series A · 2026

Better data
beats more data

Tessera is the curation engine for training sets — dedupe, score, and compose corpora that make the same model 4–9 points better on the same compute.

68
ML teams
$3.8M
ARR
$16M
raising
01 / 20
The problem

Every lab knows data quality is the highest-leverage knob — and still curates with grep and vibes.

Web-scale corpora are 40% near-duplicates and noise. Compute is priced to the penny while the thing it trains on goes unmeasured.

02 / 20
Why now
4–9 pts

benchmark improvement from curation alone at fixed compute, replicated across three model families in our published evals (2025) — cheaper than any scaling law says compute can buy.

03 / 20
Product

A refinery between your crawl and your cluster

01
Semantic dedupe
Near-duplicate detection at 100B-document scale.
02
Quality scoring
Model-based raters tuned to YOUR downstream evals.
03
Mixture optimizer
Composes domain ratios against a target eval suite.
[ SCREEN — corpus composition ]
04 / 20
Pipeline
STEP 01
Ingest
Crawls, licensed sets, synthetic pools.
STEP 02
Score
Every document rated on 14 axes.
STEP 03
Compose
Optimizer builds the training mixture.
STEP 04
Trace
Full lineage per training token.
05 / 20
Traction
01
68
ML teams
02
$3.8M
ARR
03
2.1T
tokens curated to date
04
131%
net revenue retention

Two frontier labs run Tessera inside their own VPCs — the deals that anchor the enterprise tier.

06 / 20
Growth

ARR, trailing 8 quarters

Q3 ’24Q4Q1 ’25Q2Q3Q4Q1 ’26Q2
07 / 20
Unit economics

Pricing scales with tokens, margin scales with reuse

TierPriceGross margin
Starter (100B tokens)$4k/mo68%
Scale (1T tokens)$22k/mo79%
Enterprise VPC$390k/yr84%
Blended78%
08 / 20
Market · Epoch AI 2025

Data ops: the unbudgeted 20% of every training run

0100%
14%
38%
48%
SOM — AI-native labs + model teams 14%
SAM — enterprise fine-tuning data ops 38%
TAM — training data tooling + services 48%
09 / 20
Competition
Tessera
Labeling vendors
OSS dedupe stacks
In-house scripts
Scale (tokens) → ↑ Eval-linked quality
10 / 20
Moat
01
01

Rater ensembles

Quality models distilled from 400+ customer eval suites — accuracy compounds cross-customer.

02
02

Scale engineering

100B-doc dedupe in 9 hours; rivals quote weeks.

03
03

Lineage standard

Our per-token provenance format is entering two compliance frameworks.

11 / 20
Customers

Teams training on Tessera-curated corpora

10 partners
[ logo 01 ]
[ logo 02 ]
[ logo 03 ]
[ logo 04 ]
[ logo 05 ]
[ logo 06 ]
[ logo 07 ]
[ logo 08 ]
[ logo 09 ]
[ logo 10 ]
12 / 20
Customer proof

“Same cluster, same architecture, Tessera mixture: +6.2 on our eval suite. That is a result we had budgeted eight figures of compute to reach.”

Dr. Hannah Cole
Head of Pretraining, applied-AI lab (400 GPUs)
13 / 20
Business model

Revenue per customer-year

Curation SaaS
+ Mixture optimizer
+ Lineage/compliance
Total
sources uses ending
14 / 20
plan
The plan

Own the data layer of the training stack

Compute got its platforms. Serving is getting its runtimes. Data — the highest-leverage layer — is still up for grabs.

15 / 20
Roadmap
Now
Enterprise VPC GA
From 2 design partners to 12 accounts.
H1 2027
Synthetic pipeline
Generation + verification in the same loop.
H2 2027
Multimodal curation
Image-text and audio corpora, same raters.
2028
Compliance suite
Provenance audits as regulation lands.
16 / 20
Team
[ photo ]
Dr. Felix Braun
CEO — ex-DeepMind data research
[ photo ]
Ines Costa
CTO — ex-Databricks Spark core
[ photo ]
Kenji Mori
Research — dataset papers, 8k citations
[ photo ]
Sasha Ivanova
GTM — ex-Scale AI enterprise
17 / 20
Financial plan

ARR with this round

$24M
ARR exiting 2028
25
50
75
100
$2.8M
$3.8M
$6.7M
$10.6M
$16M
$24M
H1 ’26H2 ’26H1 ’27H2 ’27H1 ’28H2 ’28
18 / 20
Use of funds

Deploying $16M over 24 months

0100%
46%
26%
28%
Research + engineering 46%
Enterprise GTM 26%
Compute + buffer 28%
19 / 20
Tessera

Train on signal.

felix@tessera.dev · tessera.dev
[ QR ]
20 / 20