Local meeting notes, benchmarked
This page is generated from the same benchmark results that generate the model catalog shipped inside the app — the numbers you read are the numbers the app acts on.
Abschrift writes meeting notes on your Mac by default: a built-in runtime plus an open model picked for your hardware, downloaded once. That promise is only honest if the local model is actually good. So we measure it — with a benchmark built around the one job that matters: turning a messy meeting transcript into notes you can act on. This edition covers 17 open models from 2B to 122B parameters, from an 8 GB MacBook Air to a 128 GB Mac Studio.
What we measure
Six synthetic meetings (no real people or companies), written to be hard in the ways real meetings are hard: short stand-ups to two-hour strategy workshops, English and German, 4–23k tokens of transcript. Each comes with a scripted ground truth: every fact, decision, and action item that correct notes must carry, with owners and deadlines. Each also contains noise that good notes must drop, like small talk and venting. And each sets attribution traps: heated accusations whose substance belongs in the notes but whose finger-pointing does not.
Every model gets the app's actual notes prompt and identical sampling. Models up to 12B ran on the app's actual bundled runtime (a pinned llama-server build, Q4-class quantization) on a base M1, so speed is measured for real. Models too big for the measurement machine were generated through the OpenRouter API at provider precision and are marked “API” below. Their Apple-Silicon speeds are estimates. The methodology section at the end explains how.
A frontier model, Claude Fable 5, grades each output against the ground truth and the full transcript: fact coverage, false claims weighted by severity, action-item fidelity, section structure, attribution neutrality, noise rejection, and whether the Discussion section alone could power follow-up work. Two grading passes per output, averaged. The same rubric grades notes written by Claude Fable 5 itself. That's the bar, and it scores 97.4.
Does a bigger model write better notes?
Results
| Model | Params | Composite | Long meetings | Facts | Correctness | Action items | Quality via | Verdict |
|---|---|---|---|---|---|---|---|---|
| Frontier baseline (Claude Fable 5) | — | 97.4 | 98.6 | 98.2 | 98.3 | 97.3 | — | The bar |
| Muse Glimmer 30B | 30B | 91.5 | 86.6 | 92.9 | 100.0 | 87.3 | API | Excellent |
| Qwen3.8 27B | 27B | 89.9 | 94.3 | 95.8 | 95.0 | 76.5 | API | Excellent |
| Qwen3.5 27B | 27B | 89.3 | 90.1 | 89.0 | 100.0 | 81.5 | API | Excellent |
| Qwen3.5 122B A10B (MoE) | 122B (A10B) | 88.8 | 88.2 | 91.2 | 97.5 | 78.8 | API | Excellent |
| GLM-4.5 Air 106B A12B (MoE) | 106B (A12B) | 88.5 | 81.2 | 84.1 | 97.5 | 87.1 | API | Excellent |
| gpt-oss 120B (MoE) | 117B (A5B) | 86.8 | 84.4 | 84.1 | 93.3 | 83.7 | API | Excellent |
| Qwen3.5 35B A3B (MoE) | 35B (A3B) | 86.6 | 85.8 | 88.1 | 97.5 | 75.2 | API | Excellent |
| Qwen3.5 9B | 9B | 86.1 | 80.9 | 86.1 | 96.7 | 69.9 | on-device | Excellent |
| Ministral 3 14B | 14B | 84.5 | 83.8 | 88.6 | 82.9 | 65.1 | API | Good |
| Qwen3.5 4B | 4B | 77.3 | 65.0 | 73.7 | 75.0 | 83.8 | on-device | Good |
| Gemma 4 E4B | 4B | 76.3 | 75.0 | 56.7 | 85.0 | 91.1 | on-device | Good |
| Gemma 4 26B A4B (MoE) | 26B (A4B) | 76.0 | 66.8 | 55.2 | 82.9 | 95.0 | API | Good |
| Gemma 4 31B | 31B | 75.8 | 65.3 | 62.2 | 82.5 | 76.7 | API | Good |
| Gemma 4 12B | 12B | 73.0 | 67.1 | 54.0 | 75.4 | 90.3 | on-device | Acceptable |
| gpt-oss 20B (MoE) | 21B (A4B) | 72.2 | 78.8 | 64.5 | 84.6 | 52.1 | API | Acceptable |
| Granite 4.1 8B | 8B | 60.5 | 50.0 | 39.9 | 67.5 | 55.9 | on-device | Not recommended |
| Qwen3.5 2B | 2B | 48.6 | 50.6 | 42.7 | 10.8 | 52.4 | on-device | Not recommended |
Scores are 0–100 composites (weights: facts 30%, correctness 20%, action items 15%, discussion quality 15%, attribution 10%, structure 5%, noise 5%). Long meetings is the composite over the 90–110-minute transcripts only — where small models fall behind first. (A…B) marks mixture-of-experts models: total (active) parameters.
Quality × speed: the map that picks your model
What your Mac gets
The app maps unified memory to a default model. It only asks you to choose when the benchmark shows a genuine quality/speed trade-off on your tier:
| Unified memory | Default model | Download | Composite | est. tok/s M1 / M2 / M4 / M4 Pro / M4 Max | Alternatives |
|---|---|---|---|---|---|
| 8 GB+ | Qwen3.5 4B | 2.74 GB | 77.3 | 14.0 / 21.6 / — / — / — | — |
| 16 GB+ | Qwen3.5 9B | 5.68 GB | 86.1 | 9.4 / 14.5 / 16.0 / — / — | Qwen3.5 4B (speed), Gemma 4 E4B (speed) |
| 32 GB+ | Qwen3.8 27B | 16.46 GB | 89.9 | — / — / 5.6 / 11.8 / — | Qwen3.5 9B (speed), Ministral 3 14B (speed) |
| 36 GB+ | Qwen3.8 27B | 16.46 GB | 89.9 | — / — / — / 11.8 / 19.3 | Qwen3.5 35B A3B (MoE) (speed), Qwen3.5 9B (speed) |
| 96 GB+ | Qwen3.8 27B | 16.46 GB | 89.9 | — / — / — / — / 19.3 | gpt-oss 120B (MoE) (speed), Qwen3.5 35B A3B (MoE) (speed) |
| 128 GB+ | Qwen3.8 27B | 16.46 GB | 89.9 | — / — / — / — / 19.3 | Qwen3.5 122B A10B (MoE) (speed), gpt-oss 120B (MoE) (speed) |
Blank cells mean the chip never shipped with that much RAM. Models above ~50 GB download in multiple parts. The app handles that automatically, each part checksum-pinned.
Why isn't Muse Glimmer 30B (composite 91.5) offered in the app? Its on-device runtime path has not been validated on real hardware yet and its mandatory reasoning makes notes take ~2.4x as long as the tier default's — it stays benchmark-only until that changes.
How long does a meeting take?
End-to-end notes generation for a ~45-minute meeting:
| Model | Notes for a 45-min meeting |
|---|---|
| Muse Glimmer 30B | ≈23.2 min (estimated incl. mandatory reasoning, slowest capable Mac: M4) |
| Qwen3.8 27B | ≈12.2 min (estimated, slowest capable Mac: M4) |
| Qwen3.5 27B | ≈10.2 min (estimated, slowest capable Mac: M4) |
| Qwen3.5 122B A10B (MoE) | ≈1.6 min (estimated, slowest capable Mac: M4 Max) |
| GLM-4.5 Air 106B A12B (MoE) | ≈1.8 min (estimated, slowest capable Mac: M4 Max) |
| gpt-oss 120B (MoE) | ≈1.4 min (estimated incl. mandatory reasoning, slowest capable Mac: M2 Max) |
| Qwen3.5 35B A3B (MoE) | ≈0.7 min (estimated, slowest capable Mac: M3 Max) |
| Qwen3.5 9B | 5.2 min (measured, M1) |
| Ministral 3 14B | ≈8.0 min (estimated, slowest capable Mac: M3) |
| Qwen3.5 4B | 3.2 min (measured, M1) |
| Gemma 4 E4B | 2.9 min (measured, M1) |
| Gemma 4 26B A4B (MoE) | ≈1.5 min (estimated, slowest capable Mac: M3) |
| Gemma 4 31B | ≈6.2 min (estimated, slowest capable Mac: M4) |
| Gemma 4 12B | 7.0 min (measured, M1) |
| gpt-oss 20B (MoE) | ≈10.2 min (estimated incl. mandatory reasoning, slowest capable Mac: M3) |
| Granite 4.1 8B | 5.0 min (measured, M1) |
| Qwen3.5 2B | 1.3 min (measured, M1) |
Methodology & honesty
- Two measurement paths. Models up to 12B were generated on the exact runtime and quantization the app ships, on a base M1 (16 GB) — quality and speed both measured. Larger models were generated via OpenRouter at the serving provider's precision (typically FP8/BF16) and judged identically. The shipped Q4-class quantization typically costs large models a point or two of quality, so treat their scores as a slight upper bound.
- Speed estimates are anchored, not guessed. Per-chip scaling comes from the community-measured llama.cpp Apple-Silicon benchmark table (a raw memory-bandwidth ratio would overestimate Ultra-class chips almost 2×). For MoE models the effective bytes read per token are calibrated against measured gpt-oss numbers — the calibration reproduces the measured anchors within ~5%, and we still publish ±25% (dense) / ±40% (MoE) error bars.
- RAM floors are computed conservatively: model file + a 32k-token KV cache + compute buffers must fit macOS's GPU working-set limit (~75% of unified memory). Floors for models we could run were validated on real hardware. The rule correctly predicts both the fits and the out-of-memory failures we observed.
- Exact models and modes. Baseline notes: Claude Fable 5, default thinking configuration, the app's exact prompt. Judging runs against the Anthropic API: each output is graded against the meeting's ground-truth fact sheet plus the full transcript. Of the 216 grading passes: 134 by Claude Opus 4.8, 82 by Claude Fable 5 — the judge model has changed over the benchmark's life. Every verdict records its judge, and the two passes per output are averaged, so they can come from different judges. Open models ran with thinking disabled (the app's own runtime setting — hybrid-thinking models like Qwen3.5 are graded in the non-thinking mode they actually ship in), except gpt-oss 20B/120B and Muse Glimmer 30B, whose formats cannot disable reasoning: those ran at default effort with reasoning tokens excluded from the graded output.
- The corpus is synthetic — that's what makes it publishable and exactly gradable. We additionally sanity-check winners on private real meetings before they enter the catalog. Aggregate impressions only: that data never leaves the machine.
- An LLM judge with ground-truth fact sheets is strict but not perfect. Two passes are averaged to reduce variance. The judge and the baseline writer are both Claude models — a judge grading its own kind is a known bias, which is why the baseline is a bar, not a contestant.
- When a new model generation looks promising, the benchmark re-runs, and this page and the in-app catalog regenerate together. A cloud provider with your own API key remains the quality ceiling — one click away in Settings.