BetterBench 0.4.0 · 05 Sep 2026 · 22:08

mlx-community/Qwen3.8-27B-oQ6

mlx-community/Qwen3.8-27B-oQ6corpus v1.05 passes/cattemp 0.7cold prefix cache (nonce)phases: decode
Self-reported — not verified by launch80

Read these first

Decode throughput by categoryMedian per-pass decode t/s at batch = 1, 5 passes per category. The dashed line is the weighted combined score.
0.01.93.75.67.4DECODE T/S (MEDIAN)chatcombined 7.3chat decode t/s: 7.3 t/s ±IQR: 0.7 CV: 7.3% passes: 5

Hover a bar for its IQR and coefficient of variation — a high CV means the category's passes disagree, so read small differences there with care.

Inter-token latency range by categoryEach bar spans the 1% low to the 99% high instantaneous token rate, with a tick at the median. Wide bars stutter; narrow bars feel smooth.
1% low → 99% highmedian
0.04,5009,00013,50018,000INSTANTANEOUS T/Schatchat 99% high: 17,376.7 t/s median: 73.5 t/s 1% low: 2.5 t/s spread: 17,374.2 t/s
Full numbers
Single-stream, batch = 1
categorypassesTTFT p50TTFT p99ITL 1% lowITL medITL 99% highdecode med±IQRCV
chat51,641.21,751.02.573.517,376.77.30.77.3%

Generated by BetterBench 0.4.0 from a corpus v1.0 run. Corpus hash 706e6cf4b5165492. Combined-score weights — code 0.3, reasoning 0.2, prose 0.15, json 0.15, file_edit 0.1, summarization 0.1. Results are only comparable within a corpus version. Stopped at max_tokens: 3/5 runs (60%) — on a thinking model a truncated run measures the thinking phase, not a complete answer. 1 percentile is marked † — it rests on fewer samples than n · tail ≥ 5 requires (a p99 needs 500 observations), so read it as roughly the worst observed rather than as a percentile. The full list is under sample_gate in the results JSON. See METHODOLOGY.md §sample-size.