Positron AI has raised $875 million at a $5 billion post-money valuation, which is a lot of money for a company whose flagship product has not yet been through a fab. The chip, called Asimov, tapes out at the end of 2026. Production silicon is expected in the second half of 2027.

Seven months ago Positron was valued at roughly $1 billion. The company is now worth five times that on the strength of simulation results and a thesis about memory.

The Structure of the Round

The $875 million is actually two instruments stacked: a $375 million Series C priced at a $3.5 billion pre-money valuation, plus up to $500 million in a Series C-1 led by NEA and Jim Clark’s investment office.

The investor list is a cross-section of everyone with a reason to want an Nvidia alternative to exist:

  • NEA and the Jim Clark Office
  • Andra Capital, Atreides Management, Valor Equity Partners
  • SemiAnalysis Capital
  • Qatar Investment Authority
  • Cisco Investments
  • Hudson River Trading
  • Naver Ventures

Four new board members came with the money: Forest Baskett of NEA, Gavin Baker of Atreides, Thomas Jermoluk of the Jim Clark Office, and Dylan Patel of SemiAnalysis. The presence of a high-frequency trading firm and a Korean internet giant on the same cap table tells you the customer profile Positron is chasing: latency-obsessed operators who run inference at volume and feel every cent of cost per token.

The Bet: Skip HBM Entirely

Asimov’s defining choice is what it does not use. There is no high-bandwidth memory. Instead, configurations run from 288GB up to 2,304GB of LPDDR5X, on a TSMC N3P process, at roughly 400 watts. Systems pack four or eight chips into a unit Positron calls Titan.

On paper this looks backwards. HBM delivers far more raw bandwidth per pin than LPDDR5X, which is why every serious AI accelerator uses it. Positron’s argument is that raw bandwidth is the wrong metric because GPUs cannot actually use most of what they have. The company claims its architecture sustains more than 90% of available memory bandwidth in practice, against under 30% for GPUs on inference workloads.

If that holds, a chip with less theoretical bandwidth but far better utilization wins on effective throughput, and wins enormously on cost, because HBM is expensive, supply-constrained and power-hungry. Two point three terabytes of LPDDR5X on a single accelerator also means very large models fit in memory without sharding across nodes, which removes an entire class of interconnect overhead.

The Claims, and Who Produced Them

Positron says Asimov delivers 26 times more tokens per dollar than Nvidia’s GB300 at peak speeds, and 2.4 times at moderate throughput. Those are striking figures and they deserve a careful reading.

Every Asimov number comes from cycle-accurate simulation. The chip has not taped out, so there is no silicon to measure. Simulation is a legitimate engineering tool and cycle-accurate models are far better than napkin math, but the gap between simulated and measured performance has humbled a long line of accelerator startups.

The comparison data for Nvidia came from SemiAnalysis InferenceX. SemiAnalysis Capital co-led this funding round. That is a real conflict and it does not make the numbers wrong, but anyone evaluating the 26x claim should notice that the benchmark provider has an equity position in the outcome.

What Positron Actually Ships Today

The company is not purely pre-revenue. Its first-generation system, Atlas, is deployed and running. Atlas is built on Intel Agilex 7 M-series FPGAs paired with HBM and DDR5, and more than 50 racks of it are in production at Oracle Cloud Infrastructure. Named customers include Parasail, Jump Trading and i3d.net.

Shipping FPGA-based systems while designing custom silicon is a sensible path. It proves the software stack, the compiler and the serving layer against real workloads, so that when Asimov arrives the only new variable is the hardware. The companies that have failed in this category usually failed on software, not silicon.

The Schedule Already Slipped

Production was previously targeted for early 2027 and is now guided to the second half. A two-quarter slip on a chip that has not taped out is unremarkable as these things go, but it is worth tracking, because the competitive window is the entire investment thesis.

By late 2027 Nvidia will have shipped another generation, AMD’s Instinct line will have advanced, and Google, Amazon and Microsoft will all have newer in-house inference silicon in their own fleets. Asimov is not being measured against today’s GB300. It is being measured against whatever exists when it arrives.

What This Means

Positron is a clean expression of the most interesting argument in AI hardware right now: that inference is a fundamentally different problem from training, and that the chips built for training are a poor fit for it. Inference is memory-bound, latency-sensitive and cost-obsessed. A design that optimizes for memory capacity and utilization rather than peak FLOPS is a coherent answer to that problem.

The thesis is sound. The execution risk is substantial. Asimov has to tape out cleanly, hit its simulated numbers in silicon, arrive on a schedule that has already moved once, and beat products that will themselves be a generation newer. The software stack has to be good enough that customers will port workloads off CUDA, which remains the hardest part of competing with Nvidia and the part that money solves least well.

The $5 billion valuation prices in most of that going right. Investors who wrote these cheques are betting that inference demand will be so large, and the cost pressure so intense, that a credible non-Nvidia option gets bought in volume almost regardless of how it stacks up on paper. On current trends, that is not a crazy bet. It is just an expensive one to be wrong about.