Skip to main content

OLIX Computing

The token factory for frontier AI

The newest models reach their best answers by spending vastly more tokens, and that demand is compounding faster than the infrastructure beneath it can scale. OLIX rebuilds the datacentre the way every great factory has been built: as a production line of specialised machines, with movement between them fast and cheap enough to finally be worth it.

TOKEN PRODUCTION LINEPREFILLSTAGE 01MEMORYSTAGE 02DECODESTAGE 03TOKENS OUT
Series B raised to build the memory layer for inference
$312MSeries B raised to build the memory layer for inference
Engineering sites across the UK, US and Canada
5Engineering sites across the UK, US and Canada
Open roles in optics, ASIC design and systems
21Open roles in optics, ASIC design and systems

The constraint

Producing a token is not one operation. It is many.

Each stage of token production places a different demand on hardware, yet every AI system in production runs all of them on the same general-purpose chip. Moving work between chips has always cost too much energy and too much time for anything else to be viable.

That single compromise is why the current approach cannot deliver high throughput and high interactivity at low cost, however large the chip grows, and why tokens stay scarce and expensive.

Throughput without the interactivity tax

Batch harder and every user waits longer. Serving each stage on hardware suited to it removes the trade-off instead of tuning around it.

A memory layer built for inference

Attention state is a memory problem wearing a compute costume. We treat it as its own stage, with its own machine, rather than renting it space on a matrix engine.

Cost per token that falls as you scale

Specialisation compounds. Every stage moved onto the right silicon takes cost out of the line, and the saving holds at rack and datacentre scale.

Energy spent on answers, not on shuttling data

Optical movement between stages cuts the interconnect overhead that general-purpose fleets pay on every single token they produce.

Headroom for reasoning-heavy models

Long chains of thought multiply token counts. A production line absorbs that growth by adding stages, not by waiting for a bigger chip.

Who it is for

  • Frontier labs training and serving reasoning models
  • Cloud and inference providers selling tokens at scale
  • Enterprises running AI workloads where latency is the product

How it works

Three moves turn a fleet of general-purpose chips into a line

  1. 01Workload profiling

    Separate the stages

    We break token production into its real constituent parts rather than treating inference as one monolithic operation. Prefill, attention, decode and memory each get measured on their own terms.

  2. 02Stage-matched silicon

    Specialise the machine

    Each stage is mapped onto hardware built for that stage alone. A memory-bound step stops competing for a matrix engine it was never going to use well.

  3. 03Photonic interconnect

    Move work at optical speed

    The line only works if the handoffs are close to free. Our optical interconnect makes movement between specialised machines cheap enough that splitting the work finally pays.

Capabilities

What sits on the line

Every part of the system exists to make one thing true: that splitting token production across specialised machines costs less than keeping it on one.

  • 01

    Stage-specialised compute

    Purpose-built machines for each phase of token production, instead of one general-purpose chip asked to do every job adequately.

  • 02

    Optical interconnect

    Laser-based movement between stages, engineered so that the energy and latency cost of a handoff stops dominating the design.

  • 03

    The memory layer for inference

    A dedicated tier for the state that inference actually spends its time on, sized and scaled independently of raw compute.

  • 04

    Production-line scheduling

    Work is routed stage to stage the way a factory routes parts, keeping every machine on the line busy with the work it is best at.

  • 05

    Datacentre-scale deployment

    Designed from the rack outward, so the architecture holds its economics at the scale frontier labs and cloud providers actually run.

  • 06

    Per-stage instrumentation

    Telemetry at every step of the line, so cost, latency and utilisation are attributable to a stage rather than averaged across a fleet.

From the field

The people running inference already know where it hurts

We had spent two years trading interactivity against throughput and calling it tuning. Splitting the stages onto hardware that suits them was the first change that moved both numbers in the same direction.

Head of InferenceFrontier AI Lab

Our cost per token had stopped improving no matter how much capacity we added. Treating the memory stage as its own machine is what finally broke that plateau for us.

Infrastructure DirectorCloud Provider

Reasoning workloads quietly tripled our token counts in a quarter. A line we can extend a stage at a time is a far easier thing to plan around than waiting on the next chip generation.

VP of EngineeringEnterprise AI Platform

Illustrative accounts from the roles OLIX works with. Names and organisations are withheld.

Work with us

Tokens should not be the scarce part of intelligence

If you are serving frontier models at scale, or building the systems that will, we would like to hear what your line looks like today.

Contact

Tell us what your line looks like

Send us the shape of your workload and we will come back with where the stages are costing you most. Every message is read by the team, not a queue.

  • Email uspress@olix.com
  • Where we buildLondon / Bristol / Austin / San Francisco / Toronto