Tag: JAX

JAX Machine Learning TPU July 13, 2026

Pallas Kernel + vLLM on TPU: Leverage Gemma Throughput to 19,511 Tokens per Second - Throughput, latency, energy, and cost results for Gemma 2B served with vLLM, with observations on the JAX and Pallas compilation stack.

JAX Machine Learning TPU July 6, 2026

From 1,540 to 15,338 Tokens per Second on a Single TPU Chip - This study significantly optimized large language model inference on a single Google TPU v5e chip, increasing throughput from 1,540 to 15,338 tokens per second. This tenfold gain was achieved through meticulous measurement and refinement of software configurations, without altering the model or hardware. The findings demonstrate substantial improvements in cost and energy efficiency, highlighting the potential of moving beyond default serving setups.

AI JAX Official Blog TensorFlow May 11, 2026

Pioneering AI-assisted code migration: How Google achieved 6x faster migration from TensorFlow to JAX - Google achieved 6x faster migration from TensorFlow to JAX. Learn how we deployed specialized, multi-agent AI systems to migrate some of Google’s largest-scale production models from TF to JAX.

 

Latest Issues




Contact

Zdenko Hrček
Třebanická 183
Prague, Czech Republic
Phone: +420 777 283 075
Email: [email protected]