Voice / results.md
Wiself's picture
small tiny fixes
271bd15
|
Raw History Blame Contribute Delete
3.47 kB

Results: two files, one tensor apart

Date: 2026-09-04. Two runs of the same benchmarks against two model files at temperature 0.0. Method: benchmarking.md (harness, settings, adaptations, extraction rules — everything needed to reproduce).

Setup

  • File A: TheDrummer_Orion-26B-A4B-v1-IQ4_NL.gguf (Gemma 4 26B MoE, IQ4), served via llama-server, unmodified.
  • File B: lm_head-TheDrummer_Orion-26B-A4B-v1-IQ4_NL.gguf — file A with its lm_head tensor replaced via voice cast, Q8_0. Nothing else changed.
  • Both files: temperature 0.0, seed 1234, identical prompts, identical token budgets (ARC 1024, MMLU 2048, GSM8K 1024), identical answer extraction. Unanswered items got one rescue pass at 4096 (MMLU) / 2048 (GSM8K), applied to both files symmetrically. Rescue tie-break: on split pairs (two verdicts for one doc) the decided verdict beats unanswered.

Scores

Task n A correct B correct A acc B acc
ARC-Challenge (generative) 100 94 95 0.940 0.950
MMLU (generative, 2 × 57 subjects) 114 101 103 0.886 0.904
GSM8K 100 76 92 0.760 0.920

Observations

  1. ARC (+1) and MMLU (+2) differences fall within slice noise (±5 points at n≈100). GSM8K (+16) exceeds it; see point 2.
  2. The GSM8K difference sits in unanswered counts. File A left 23 items silent (empty content after thought exhausted the budget); file B left 6. On items both files answered, accuracy is 0.987 (A) vs 0.979 (B). The traces hold whatever else there is to see.
  3. Thinking lengths roughly match. No pairing script was archived, so no median is claimed; per-file thinking lengths sit in the same range.
  4. Error overlap is minimal. One GSM8K item was wrong under both files; remaining errors are different items per file.

What this does not show

  • Small slices: ~100 items per task. Differences under ~5 points are noise.
  • Generative scoring is not comparable to published logprob leaderboards; compare only file A vs. file B within this table.
  • One model pair, one replaced tensor, greedy decoding only. Other pairs may differ — that is what the harness in benchmarking/ is for.
  • Writing voice is measured separately below — same two files, same neutrality, no judge model.

Evidence

Per file, archived under benchmarking/runs/file-a/ and benchmarking/runs/file-b/: harness results_*.json and samples_*.jsonl for every run, rescue files for the second passes, and the thinking log (every prompt, answer, and reasoning trace). No personal information in any of it — server file paths are redacted to run labels at the log source.

Part II — writing voice: 32 prompts, two files

Method: creative-bench.md. 32 EQ-Bench v3 prompts, one generation per prompt per file at temperature 0.7, 6000-token budget. Measured, not judged: shared trigram vocabulary, cliche density, texture stats. Blind A/B pairs archived for readers.

Measure File A File B
Shared trigram vocab, mean (max) 0.013 (0.027) across all 32 same pair
Cliche hits per 100 words 0.141 0.075
Avg words per piece 1018.0 878.6
Lexical diversity (TTR) 0.4439 0.5054
Avg sentence words 9.3 8.8

Per-prompt exhibits: creative-bench/runs/file-a-creative.jsonl and file-b-creative.jsonl (prompt, full text, full reasoning trace each).