OpenAI announced that its first-generation custom inference silicon, named Jalapeño, delivered record mixed-token throughput per kilowatt on the public InferenceX benchmark, outpacing the best commercial accelerators currently in use. The chip posted 53.7× the throughput-per-energy of the prior best system while also shaving latency across GPT-OSS 120B, DeepSeek R1 and Kimi K2 model families. This concrete performance leap is the newest data point in OpenAI’s broader compute strategy, which treats the entire AI stack—from data-center design to developer APIs—as a single leverable system.

Jalapeño inference chip and stack integration

  • Throughput per kW: 53.7× higher than the previous best accelerator (InferenceX results).
  • Token latency: Lower across all three tested models, translating to faster user-facing responses.
  • Mixed-token throughput: 535.28 tokens/s/user for GPT-OSS, 169.41 tokens/s/user for DeepSeek R1, 182.46 tokens/s/user for Kimi K2.
  • Energy draw: Measured in mixed-tokens per second per kilowatt, confirming a clear advantage in power-constrained environments.

These figures prove that purpose-built silicon can beat off-the-shelf GPUs and ASICs when the software stack is co-designed. OpenAI now controls the full stack—model, serving software, memory hierarchy, network fabric, and chip—allowing simultaneous optimization across all layers.

System-wide leverage: breadth, depth, and choice

OpenAI’s compute portfolio still includes Microsoft Azure, NVIDIA GPUs, AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy and SoftBank. By keeping a diversified set of partners, the company can allocate each workload to the hardware that offers the best capability-cost ratio. For high-throughput inference, Jalapeño becomes the default; for frontier training, NVIDIA’s H100-class GPUs remain dominant; for low-latency edge agents, specialized ASICs from partners fill the niche.

  • Pareto frontier focus: OpenAI continuously seeks the optimal mix of speed, reliability, efficiency and cost for each workload.
  • Economic discipline: Premium hardware is reserved for tasks where raw capability matters (e.g., large-scale model training), while Jalapeño and other efficiency-focused silicon drive down the cost of serving billions of daily queries.
  • Local data-center leverage: Project Camellia in Georgia illustrates how OpenAI designs facilities around workload patterns, integrating closed-loop water cooling and renewable energy contracts to further improve the cost per token.

The Jevons paradox in AI

OpenAI’s internal Artificial Analysis Coding Agent Index shows GPT-5.6 Sol achieving a new high on reasoning tasks while using 54% fewer output tokens than competing models. The direct impact for customers includes faster completion of complex queries, reduced retry rates and lower total cost of ownership, and the ability to run longer, more intricate workflows without hitting cost ceilings. When each token becomes cheaper, new use-cases emerge—real-time contract analysis, live financial scenario modeling, and continuous code testing—all of which were previously cost-prohibitive. This mirrors the Jevons paradox: greater efficiency expands overall consumption, driving fresh revenue streams for enterprises that embed AI agents into core processes.

Compounding advantage: reinvestment loop

Higher efficiency frees capital that OpenAI redirects into research, safety, and further infrastructure upgrades. The cycle is self-reinforcing: better hardware yields cheaper compute, which fuels more model training, which in turn produces more capable models that demand even more compute. OpenAI’s ability to iterate on both silicon and software in lockstep creates a competitive moat that is difficult for rivals to replicate without comparable vertical integration.

New analysis: incentives, risks, and next steps

OpenAI’s incentive to push silicon efficiency is twofold. First, lower token costs directly improve margins on its API business, which accounts for a growing share of revenue. Second, by owning the silicon layer, OpenAI can dictate performance baselines that third-party cloud providers must match, reducing the risk of price wars. However, this strategy carries risks. Concentrating design expertise in a single organization raises supply-chain vulnerability; a fabrication defect or a delay at the foundry could ripple through the entire service stack. Moreover, tighter integration may limit openness, making it harder for external developers to benchmark or fine-tune models on alternative hardware.

From a market perspective, the efficiency gains could accelerate adoption in cost-sensitive sectors such as fintech, healthcare, and education. Companies that previously avoided large language models due to token expense may now experiment with real-time assistants, expanding the overall AI market size. At the same time, regulators may scrutinize the environmental impact claims, demanding transparent reporting of energy savings versus total data-center consumption.

What developers should watch next

  • Silicon roadmap: OpenAI has hinted at next-generation Jalapeño chips that will push throughput per kW even higher and add on-chip memory optimizations for transformer-style workloads.
  • API pricing adjustments: As serving costs decline, OpenAI may revise its usage tiers, potentially lowering the barrier for startups to adopt high-throughput models.
  • Ecosystem impact: Third-party cloud providers will need to decide whether to integrate Jalapeño-based instances or risk losing high-volume inference customers.
  • Open model weights: The broader community can benchmark against OpenAI’s claims using publicly available model checkpoints, such as those hosted on the open model weights repository.

The Jalapeño announcement marks a decisive shift from reliance on external accelerators to a more autonomous, performance-driven compute stack. By aligning hardware, software, and data-center design, OpenAI not only tightens its cost structure but also sets a new benchmark for what integrated AI infrastructure can achieve.


Related coverage

Explore more on this topic