# Results: two files, one tensor apart Date: 2026-09-04. Two runs of the same benchmarks against two model files at temperature 0.0. Method: [`benchmarking.md`](./benchmarking.md) (harness, settings, adaptations, extraction rules — everything needed to reproduce). ## Setup - File A: `TheDrummer_Orion-26B-A4B-v1-IQ4_NL.gguf` (Gemma 4 26B MoE, IQ4), served via llama-server, unmodified. - File B: `lm_head-TheDrummer_Orion-26B-A4B-v1-IQ4_NL.gguf` — file A with its `lm_head` tensor replaced via `voice cast`, Q8_0. Nothing else changed. - Both files: temperature 0.0, seed 1234, identical prompts, identical token budgets (ARC 1024, MMLU 2048, GSM8K 1024), identical answer extraction. Unanswered items got one rescue pass at 4096 (MMLU) / 2048 (GSM8K), applied to both files symmetrically. Rescue tie-break: on split pairs (two verdicts for one doc) the decided verdict beats `unanswered`. ## Scores | Task | n | A correct | B correct | A acc | B acc | |---|---|---|---|---|---| | ARC-Challenge (generative) | 100 | 94 | 95 | 0.940 | 0.950 | | MMLU (generative, 2 × 57 subjects) | 114 | 101 | 103 | 0.886 | 0.904 | | GSM8K | 100 | 76 | 92 | 0.760 | 0.920 | ## Observations 1. **ARC (+1) and MMLU (+2) differences fall within slice noise** (±5 points at n≈100). GSM8K (+16) exceeds it; see point 2. 2. **The GSM8K difference sits in unanswered counts.** File A left 23 items silent (empty content after thought exhausted the budget); file B left 6. On items both files answered, accuracy is 0.987 (A) vs 0.979 (B). The traces hold whatever else there is to see. 3. **Thinking lengths roughly match.** No pairing script was archived, so no median is claimed; per-file thinking lengths sit in the same range. 4. **Error overlap is minimal.** One GSM8K item was wrong under both files; remaining errors are different items per file. ## What this does not show - Small slices: ~100 items per task. Differences under ~5 points are noise. - Generative scoring is not comparable to published logprob leaderboards; compare only file A vs. file B within this table. - One model pair, one replaced tensor, greedy decoding only. Other pairs may differ — that is what the harness in `benchmarking/` is for. - Writing voice is measured separately below — same two files, same neutrality, no judge model. ## Evidence Per file, archived under `benchmarking/runs/file-a/` and `benchmarking/runs/file-b/`: harness `results_*.json` and `samples_*.jsonl` for every run, rescue files for the second passes, and the thinking log (every prompt, answer, and reasoning trace). No personal information in any of it — server file paths are redacted to run labels at the log source. ## Part II — writing voice: 32 prompts, two files Method: [`creative-bench.md`](./creative-bench.md). 32 EQ-Bench v3 prompts, one generation per prompt per file at temperature 0.7, 6000-token budget. Measured, not judged: shared trigram vocabulary, cliche density, texture stats. Blind A/B pairs archived for readers. | Measure | File A | File B | |---|---|---| | Shared trigram vocab, mean (max) | 0.013 (0.027) across all 32 | same pair | | Cliche hits per 100 words | 0.141 | 0.075 | | Avg words per piece | 1018.0 | 878.6 | | Lexical diversity (TTR) | 0.4439 | 0.5054 | | Avg sentence words | 9.3 | 8.8 | Per-prompt exhibits: `creative-bench/runs/file-a-creative.jsonl` and `file-b-creative.jsonl` (prompt, full text, full reasoning trace each).