OLIX Computing
The token factory for frontier AI
The newest models reach their best answers by spending vastly more tokens, and that demand is compounding faster than the infrastructure beneath it can scale. OLIX rebuilds the datacentre the way every great factory has been built: as a production line of specialised machines, with movement between them fast and cheap enough to finally be worth it.
- Series B raised to build the memory layer for inference
- $312MSeries B raised to build the memory layer for inference
- Engineering sites across the UK, US and Canada
- 5Engineering sites across the UK, US and Canada
- Open roles in optics, ASIC design and systems
- 21Open roles in optics, ASIC design and systems
The constraint
Producing a token is not one operation. It is many.
Each stage of token production places a different demand on hardware, yet every AI system in production runs all of them on the same general-purpose chip. Moving work between chips has always cost too much energy and too much time for anything else to be viable.
That single compromise is why the current approach cannot deliver high throughput and high interactivity at low cost, however large the chip grows, and why tokens stay scarce and expensive.
Throughput without the interactivity tax
Batch harder and every user waits longer. Serving each stage on hardware suited to it removes the trade-off instead of tuning around it.
A memory layer built for inference
Attention state is a memory problem wearing a compute costume. We treat it as its own stage, with its own machine, rather than renting it space on a matrix engine.
Cost per token that falls as you scale
Specialisation compounds. Every stage moved onto the right silicon takes cost out of the line, and the saving holds at rack and datacentre scale.
Energy spent on answers, not on shuttling data
Optical movement between stages cuts the interconnect overhead that general-purpose fleets pay on every single token they produce.
Headroom for reasoning-heavy models
Long chains of thought multiply token counts. A production line absorbs that growth by adding stages, not by waiting for a bigger chip.
Who it is for
- Frontier labs training and serving reasoning models
- Cloud and inference providers selling tokens at scale
- Enterprises running AI workloads where latency is the product
How it works
Three moves turn a fleet of general-purpose chips into a line
- 01Workload profiling
Separate the stages
We break token production into its real constituent parts rather than treating inference as one monolithic operation. Prefill, attention, decode and memory each get measured on their own terms.
- 02Stage-matched silicon
Specialise the machine
Each stage is mapped onto hardware built for that stage alone. A memory-bound step stops competing for a matrix engine it was never going to use well.
- 03Photonic interconnect
Move work at optical speed
The line only works if the handoffs are close to free. Our optical interconnect makes movement between specialised machines cheap enough that splitting the work finally pays.
Capabilities
What sits on the line
Every part of the system exists to make one thing true: that splitting token production across specialised machines costs less than keeping it on one.
- 01
Stage-specialised compute
Purpose-built machines for each phase of token production, instead of one general-purpose chip asked to do every job adequately.
- 02
Optical interconnect
Laser-based movement between stages, engineered so that the energy and latency cost of a handoff stops dominating the design.
- 03
The memory layer for inference
A dedicated tier for the state that inference actually spends its time on, sized and scaled independently of raw compute.
- 04
Production-line scheduling
Work is routed stage to stage the way a factory routes parts, keeping every machine on the line busy with the work it is best at.
- 05
Datacentre-scale deployment
Designed from the rack outward, so the architecture holds its economics at the scale frontier labs and cloud providers actually run.
- 06
Per-stage instrumentation
Telemetry at every step of the line, so cost, latency and utilisation are attributable to a stage rather than averaged across a fleet.
From the field
The people running inference already know where it hurts
We had spent two years trading interactivity against throughput and calling it tuning. Splitting the stages onto hardware that suits them was the first change that moved both numbers in the same direction.
Our cost per token had stopped improving no matter how much capacity we added. Treating the memory stage as its own machine is what finally broke that plateau for us.
Reasoning workloads quietly tripled our token counts in a quarter. A line we can extend a stage at a time is a far easier thing to plan around than waiting on the next chip generation.
Illustrative accounts from the roles OLIX works with. Names and organisations are withheld.
Work with us
Tokens should not be the scarce part of intelligence
If you are serving frontier models at scale, or building the systems that will, we would like to hear what your line looks like today.
Contact
Tell us what your line looks like
Send us the shape of your workload and we will come back with where the stages are costing you most. Every message is read by the team, not a queue.
- Email uspress@olix.com
- Where we buildLondon / Bristol / Austin / San Francisco / Toronto