munaGemma 4 26B A4B
- compiled code size
- 96MB
- compiled code size
- artifact pull, vs 25.8GB
- 760MB
- artifact pull, vs 25.8GB
- files pulled, vs 193,735
- 5
- files pulled, vs 193,735
- CTTFT @p90 · Muna (bare metal)
- 6.3s
- CTTFT @p90 · Muna (bare metal)
- lower median CTTFT vs SGLang (Modal)
- 18×
- lower median CTTFT vs SGLang (Modal)
- TTFT @16k-token prompts
- 147ms
- TTFT @16k-token prompts
- tok/s across prompt categories
- 206–593
- tok/s across prompt categories
Every stat is computed from the raw result files at the bottom of this page. See methodology for how each metric is defined and measured.
Artifact size
Everything that must be pulled onto a machine before it can serve a single token: the stock SGLang container image versus the entire Muna-compiled Gemma 4 artifact. Hover any slice to see what it is.
lmsysorg/sglang:latest
Hover a slice to inspect a file.
Muna-compiled Gemma 4
Hover a slice to inspect a file.
Cold-start Time to First Token
Cold-start Time to First Token: the time from a cold start to a user's first token, including container/artifact fetch, runtime init, model load, prefill, and first decode. Page cache was flushed before every bare-metal run.
Time to First Token
Warm time to first token across prompt lengths at batch size 1. Kernels are sourced from FlashInfer, TensorRT-LLM, and CUTLASS at compile time.
Time to First Token vs. Prompt Length
Decode throughput
Decode throughput across reasoning, math, code, knowledge, and summarization prompts at batch size 1, with speculative decoding via DFlash.
Decode Throughput by Prompt Category
Run configuration
- Model
- @google/gemma-4-26b-a4b-it
- Weights
- google/gemma-4-26B-A4B-it
- Hardware
- NVIDIA B200
- Run date
- May 2026
- Baseline
- lmsysorg/sglang:latest
- Cold boots
- 5 per target
- Host
- RunPod B200 (kv85we196c0zw2)
- Storage
- local NVMe overlay (/), 8x SOLIDIGM SB5PH27X076T md0 RAID10
- Cache eviction
- posix_fadvise(POSIX_FADV_DONTNEED) per boot; /proc/sys/vm/drop_caches unavailable in container
- TTFT rounds
- 50 per prompt length, 3 warmup, batch size 1
- Throughput rounds
- 20 per category, 512 tokens, batch size 1
Raw data
Every chart on this page is rendered directly from these result files, including the raw per-run samples. Rerun the benchmark with the reproduction script.
Run this model
Serve Gemma 4 26B A4B on Muna through the OpenAI or Anthropic API, or deploy the same compiled model to your own compute. Get an API key.
$ muna deploy @google/gemma-4-26b-a4b-it --provider modal --gpu b200