OpenAI has unveiled the Jalapeno ASIC, a custom chip paired with its Gluon kernel programming language. Built on TSMC’s N3P node with an N3E I/O chiplet, the hardware features HBM4 memory, MXFP4 support, and a single-token prediction architecture for AI tasks.
The chip delivers 13.4 PFLOPs of MXFP4 performance and 15.4 TB/s of memory bandwidth. Utilizing weight-stationary systolic arrays and 64-bit Out-of-Order scalar cores, it maintains a 700W TDP while typically consuming 550W during active inference workloads.
Jalapeno outperforms NVIDIA GB200 and GB300 systems, which use HBM3E. On the InferenceX benchmark, it achieved 1.5x to 1.9x better efficiency per watt and 1.7x to 3.6x lower latency, marking a major shift in high-performance AI compute and power efficiency.