Lumen.ai
Series A · 2026
Series A · 2026

Serve models
at the speed of demand

Lumen is an inference runtime that packs 3.2× more tokens per GPU — cold starts in milliseconds, autoscaling that follows traffic, bills that follow usage.

240
engineering teams
$5.6M
ARR
$20M
raising
01 / 20
The problem

Inference is 90% of an AI product’s compute bill — and most GPUs serving it sit 70% idle.

Teams over-provision for peaks, eat cold starts at the valleys, and rewrite serving code for every model family. The runtime layer is a decade behind the models.

02 / 20
Why now
$97B

projected annual inference spend by 2028 — passing training spend for the first time (Epoch AI compute report, 2025). Every efficiency point is a billion-dollar line.

03 / 20
Product

One runtime. Any model. Full GPUs.

01
Continuous batching++
Speculative + paged attention tuned per architecture, automatically.
02
Millisecond cold starts
Snapshot-restore of loaded weights — scale-to-zero without penalty.
03
Bring your cloud
Runs in your VPC on your committed spend, or on ours.
[ SCREEN — GPU utilization board ]
04 / 20
How it works
STEP 01
Deploy
One manifest per model; any weights.
STEP 02
Pack
Requests bin-packed across the fleet.
STEP 03
Scale
Replicas follow p99 latency, not schedules.
STEP 04
Meter
Per-token billing to the team level.
05 / 20
Traction

Production numbers, trailing quarter

01
240
engineering teams
02
$5.6M
ARR
03
3.2×
tokens/GPU vs baseline vLLM
04
41B
tokens served / day
05
99.98%
uptime, trailing 12 mo
06
142%
net revenue retention
06 / 20
Growth

ARR, trailing 8 quarters

Q3 ’24Q4Q1 ’25Q2Q3Q4Q1 ’26Q2
07 / 20
Unit economics

Customer economics at the median

MetricValue
Median customer spend / mo$1,940
Customer GPU savings vs self-run54%
Lumen gross margin (SaaS)76%
CAC payback7 months
NRR142%
08 / 20
Market · Epoch AI 2025

Inference infrastructure: the layer everyone pays

0100%
37%
52%
SOM — AI-native product teams 11%
SAM — self-hosted inference 37%
TAM — all inference compute 52%
09 / 20
Competition
Lumen
Managed endpoints
OSS runtimes
Cloud ML platforms
GPU efficiency → ↑ Deployment flexibility
10 / 20
Moat
01
01

Scheduler telemetry

Packing decisions trained on 41B tokens/day of real traffic shapes.

02
02

Kernel library

Hand-tuned kernels for 9 GPU SKUs — 18 months of low-level work.

03
03

Switching gravity

Billing, quotas and SLOs live in the runtime; ripping it out means rebuilding ops.

11 / 20
Customers

Teams serving production traffic on Lumen

10 partners
[ logo 01 ]
[ logo 02 ]
[ logo 03 ]
[ logo 04 ]
[ logo 05 ]
[ logo 06 ]
[ logo 07 ]
[ logo 08 ]
[ logo 09 ]
[ logo 10 ]
12 / 20
Customer proof

“We cut our inference bill 58% the week we switched and deleted four thousand lines of serving glue. Lumen is the first infra bill I have ever enjoyed reading.”

Dev Sharma
CTO, legal-AI platform (Series B)
13 / 20
Business model

Revenue per customer-year

Runtime SaaS
+ Enterprise VPC
+ Priority kernels
Total
sources uses ending
14 / 20
plan
The plan

Become the default serving layer

Runtimes standardize fast and stay standard for a decade — the next 24 months decide whose manifest everyone writes.

15 / 20
Roadmap
Now
Enterprise VPC GA
SOC 2 Type II complete; 6 enterprise pilots.
H1 2027
Multi-modal serving
Image + audio pipelines, same scheduler.
H2 2027
Fleet marketplace
Burst across customers’ committed spend.
2028
Edge tier
Same manifest from H100 rack to on-prem box.
16 / 20
Team
[ photo ]
Wei Zhang
CEO — ex-NVIDIA Triton core
[ photo ]
Olga Petrov
CTO — ex-Google TPU serving
[ photo ]
Marcus Hill
Kernels — ex-Meta PyTorch
[ photo ]
Julia Reyes
GTM — ex-Datadog enterprise
17 / 20
Financial plan

ARR with this round

$34M
ARR exiting 2028
25
50
75
100
$3.8M
$5.6M
$8.8M
$14M
$22M
$34M
H1 ’26H2 ’26H1 ’27H2 ’27H1 ’28H2 ’28
18 / 20
Use of funds

Deploying $20M over 24 months

0100%
48%
28%
24%
Engineering (kernels, scheduler) 48%
Enterprise GTM 28%
Fleet capex + buffer 24%
19 / 20
Lumen.ai

Every GPU, fully lit.

wei@lumen.ai · lumen.ai
[ QR ]
20 / 20