Every commit scored for change-risk against this repo's own history, so 'elevated' means elevated here rather than on some global curve.
Needs review
25 commits sit in this repo's top risk tercile, which is 34% of the 74scored. The cut is drawn against this codebase's own history rather than a global curve, so a quiet repo still fills its top band, and here it starts at 9.4 out of 10. What pushes a commit up is size and spread together: a large change confined to one area scores below a smaller one scattered across a dozen files.
Commit categories over time, read off the subject line. Fixes carry the accent because that is the series this chart exists to show.
Consistently other-driven across its history.
Ranked by change-risk, highest first. Priority is a tercile of this repo's own distribution, so a quiet repo still fills its top band.
| # | Commit | Author | When | Lines | Risk | Top driver |
|---|---|---|---|---|---|---|
| 1 | 943b0da0fix(m5-full-sample): harden run-swe.js with retry+persistence, bring 4 short results to full n, reconcile scorecard | Travis Boudreaux | 1mo ago | +4.7K -121 | 97%Elevated | large diff (many lines added) |
| 2 | 258bdb63data(evals): run1 results + per-skill verdicts | Travis Boudreauxclaudeassisted | 1mo ago | +6.4K -0 | 97%Elevated | large diff (many lines added) |
| 3 | 7e606b31Add 20 new thinking skills and quality improvement scripts | Travis Boudreauxclaudeagent | 6mo ago | +8.0K -26 | 97%Elevated | large diff (many lines added) |
| 4 | 114a97c6Initial release: 18 thinking skills for Claude Code | Travis Boudreauxclaudeagent | 6mo ago | +5.7K -0 | 97%Elevated | large diff (many lines added) |
| 5 | af234193feat(m3-calibration-freeze): run placebo-only calibration, freeze decisive splits at 40-70% band, mark calibration-only items | Travis Boudreaux | 1mo ago | +2.4K -74 | 91%Elevated | large diff (many lines added) |
| 6 | 178e47c2feat(m3-datasets): build candidate item pools per evaluation family with manifests | Travis Boudreaux | 1mo ago | +2.8K -0 | 91%Elevated | large diff (many lines added) |
| 7 | 2d67b14ctest(evals): expand 17 pairwise behavioral sets from 3 to 25 problems each | Travis Boudreauxclaudeassisted | 1mo ago | +1.9K -0 | 91%Elevated | large diff (many lines added) |
| 8 | 8233eebddata(evals): paired-experiment + capability-ladder results | Travis Boudreauxclaudeassisted | 1mo ago | +2.6K -0 | 91%Elevated | large diff (many lines added) |
| 9 | 4980673cfeat(reviews): adversarial review + lift-committee tooling | Travis Boudreauxclaudeassisted | 1mo ago | +4.8K -0 | 91%Elevated | large diff (many lines added) |
| 10 | 80eb3ba6Add thinking-model-router as single entry point for all mental models | Travis Boudreauxclaudeagent | 6mo ago | +3.7K -3 | 91%Elevated | large diff (many lines added) |
| 11 | 5f927482feat(m3-harder-data-quantitative): source harder quantitative-uncertainty items, recalibrate, record CEILING-NEEDS-HARDER-DATA | Travis Boudreaux | 1mo ago | +1.3K -348 | 82%Elevated | large diff (many lines added) |
| 12 | 7f9b23e1feat(m2-harness): judge reliability validator + paired stats, calibration, replication, distractor scoring runners | Travis Boudreaux | 1mo ago | +1.3K -1 | 82%Elevated | large diff (many lines added) |
| 13 | d34f00f0fix(evals/contracts): correct eval_family for map-territory per VAL-CONTRACT-011 | Travis Boudreaux | 1mo ago | +922 -0 | 82%Elevated | large diff (many lines added) |
| 14 | 07137012fix(m1-scorecard): resolve 3 blocking scrutiny issues | Travis Boudreaux | 1mo ago | +1.6K -0 | 82%Elevated | large diff (many lines added) |
| 15 | 653b66acfeat(evals): authored datasets + HuggingFace ingester/classifier | Travis Boudreauxclaudeassisted | 1mo ago | +1.6K -0 | 82%Elevated | large diff (many lines added) |
| 16 | f2a7bbb0feat(evals): four-tier + objective eval runners | Travis Boudreauxclaudeassisted | 1mo ago | +895 -0 | 82%Elevated | large diff (many lines added) |
| 17 | 85a2d544feat(m3-harder-data-conceptual-systems): source harder conceptual + systems items, recalibrate, record CEILING-NEEDS-HARDER-DATA | Travis Boudreaux | 1mo ago | +1.0K -53 | 76%Elevated | large diff (many lines added) |
| 18 | ebf6dbc3feat(m3-datasets): source harder security items (diverse CWEs + near-miss safe variants) and harder routing items (NONE cases + close distractors), recalibrate both pools at K_TRIALS=5 | Travis Boudreaux | 1mo ago | +966 -1.8K | 76%Elevated | large diff (many lines added) |
| 19 | e941f9c0feat(skills): apply audit best-practices to all 39 skills | Travis Boudreauxclaudeassisted | 1mo ago | +1.3K -2.5K | 76%Elevated | large diff (many lines added) |
| 20 | d64a133bdocs(analysis): evaluation suite, audit, and elevate-or-kill synthesis | Travis Boudreauxclaudeassisted | 1mo ago | +762 -0 | 76%Elevated | large diff (many lines added) |
| 21 | a0b2554dfeat(m4): rework six high-mechanism cohort skills from pre-registered specs | Travis Boudreaux | 1mo ago | +1.0K -2.4K | 71%Elevated | large diff (many lines added) |
| 22 | 06997ea9feat(experiments): worktree orchestrator + paired/stack/capability runners | Travis Boudreauxclaudeassisted | 1mo ago | +582 -0 | 71%Elevated | large diff (many lines added) |
| 23 | 5fcfabcefeat(evals): SQLite store + self-contained HTML dashboard | Travis Boudreauxclaudeassisted | 1mo ago | +807 -0 | 71%Elevated | large diff (many lines added) |
| 24 | af491844docs(analysis): refresh eval synthesis, backup record, dashboard, quality report | Travis Boudreauxclaudeassisted | 1mo ago | +992 -405 | 68%Elevated | large diff (many lines added) |
| 25 | a3ce58d1feat(evals): isolated harness core | Travis Boudreauxclaudeassisted | 1mo ago | +507 -0 | 68%Elevated | large diff (many lines added) |
| 26 | fa57359ffeat(m6): executive synthesis, consolidation plan, stale-claim cleanup, active-pull future-work, README refresh | Travis Boudreaux | 1mo ago | +567 -4 | 66%Typical | large diff (many lines added) |
| 27 | b11f5538feat(m5): powered primary batch — 7 gated skills evaluated, all NO-LIFT | Travis Boudreaux | 1mo ago | +476 -19 | 64%Typical | large diff (many lines added) |
| 28 | 401aef70fix(m3-conceptual-systems-reconcile): re-freeze decisive splits from harder calibrated pools, reconcile manifest counts, fix source refs, set systems in_band=true | Travis Boudreaux | 1mo ago | +375 -309 | 64%Typical | large diff (many lines added) |
| 29 | 30e2ad30feat(m3-real-calibration): add batch mode to run-calibration.js, run real calibration on all 6 pools, replace heuristic baselines with measured values | Travis Boudreaux | 1mo ago | +348 -189 | 61%Typical | large diff (many lines added) |
| 30 | 093dac38feat(evals): add abstention + binary-decision runners and balanced datasets | Travis Boudreauxclaudeassisted | 1mo ago | +444 -14 | 61%Typical | large diff (many lines added) |
| 31 | 9dcb3a71eval(coverage): confirm leads collapse at power; close 7 gaps; objective coverage 13→17 | Travis Boudreauxclaudeassisted | 1mo ago | +315 -2 | 58%Typical | large diff (many lines added) |
| 32 | 9e803269Add roadmap for future thinking skills | Travis Boudreauxclaudeagent | 6mo ago | +272 -0 | 58%Typical | large diff (many lines added) |
| 33 | 0091d63bfeat(evals): close coverage gaps — 6 authored objective sets + StrategyQA + run batches | Travis Boudreauxclaudeassisted | 1mo ago | +220 -4 | 56%Typical | large diff (many lines added) |
| 34 | fcb8c41afix(m5-decisive-rerun): rerun 4 debugging skills on FROZEN 224-item decisive split, fix scorecard verdicts | Travis Boudreaux | 1mo ago | +220 -33 | 53%Typical | large diff (many lines added) |
| 35 | 0f50885dfix(m2-harness): add McNemar cc/midp aliases, fix replication ELEVATE gate (positive lift only), support pretty-JSON parsing | Travis Boudreaux | 1mo ago | +201 -19 | 53%Typical | large diff (many lines added) |
| 36 | a055d7d2feat(evals): judge rubrics + statistical power analysis | Travis Boudreauxclaudeassisted | 1mo ago | +206 -0 | 53%Typical | large diff (many lines added) |
| 37 | f5a79058fix(m3-security-nearmiss-26-audit): fix sec-nearmiss-26 buffer-write vulnerability, complete safe-label audit, update aggregate baseline | Travis Boudreaux | 1mo ago | +152 -151 | 50%Typical | large diff (many lines added) |
| 38 | cb2b0515fix(m2-harness): calibration k>1 trials per item, distractor scoring on real dataset | Travis Boudreaux | 1mo ago | +161 -86 | 50%Typical | large diff (many lines added) |
| 39 | e0bb0d91feat(m4): add trigger cards to 17 remaining contract-identified skills for full trigger-vs-full coverage | Travis Boudreaux | 1mo ago | +171 -0 | 47%Typical | large diff (many lines added) |
| 40 | 2f6f7741eval(wave-c): record 6 powered verdicts; headline — debugging ELEVATEs do not replicate | Travis Boudreauxclaudeassisted | 1mo ago | +185 -3 | 47%Typical | large diff (many lines added) |
| 41 | df34adadfeat(m4): add trigger cards and strengthen redirects for 9 quarantine/kill candidates | Travis Boudreaux | 1mo ago | +197 -96 | 45%Typical | large diff (many lines added) |
| 42 | 4a176d9ffeat(scientific-method): ship hypothesis-differential debugging, retire v2 prototype | Travis Boudreauxclaudeassisted | 1mo ago | +117 -104 | 44%Typical | large diff (many lines added) |
| 43 | 596a3b1dfix(m6-synthesis): resolve 5 blocking consistency issues (M6 scrutiny round-1) | Travis Boudreaux | 1mo ago | +93 -56 | 42%Typical | large diff (many lines added) |
| 44 | 291728d9chore(planning): file-based plan, findings, and session log | Travis Boudreauxclaudeassisted | 1mo ago | +115 -0 | 42%Typical | large diff (many lines added) |
| 45 | 3fe659f4docs: adversarial accuracy pass + world-class README rewrite | Travis Boudreaux | 1mo ago | +128 -50 | 39%Typical | large diff (many lines added) |
| 46 | 0c97d3b6fix(m3-datasets): fix fault-localization scorer parsing + truncation bugs, recalibrate debugging family | Travis Boudreaux | 1mo ago | +87 -37 | 39%Typical | large diff (many lines added) |
| 47 | 2793a823fix(m6-synthesis): exhaustive [SUPERSEDED] tag sweep across ELEVATE-OR-KILL.md | Travis Boudreaux | 1mo ago | +75 -75 | 36%Typical | large diff (many lines added) |
| 48 | 004d09b8feat(m5): add fresh scientific-method replication (+8.0pp, p=0.001) and scorecard updates | Travis Boudreaux | 1mo ago | +74 -21 | 36%Typical | large diff (many lines added) |
| 49 | 5bc1dc31fix(m6-synthesis): add QUARANTINE-REDIRECT disposition to scorecard and tag all eval claims with provenance tuples | Travis Boudreaux | 1mo ago | +56 -9 | 34%Typical | large diff (many lines added) |
| 50 | af0c31e8docs(repo): eval-informed metadata, README banner, agent guidelines | Travis Boudreauxclaudeassisted | 1mo ago | +91 -5 | 33%Below typical | large diff (many lines added) |
Two views of the same model: where the cuts fall, and what commit shape lands you above them.
Every scored commit, binned on the raw 0 to 10 score rather than the percentile. Percentile ranks are uniform by construction, so that axis has no shape to draw. The dashed lines are the tercile cuts behind each row's priority pill.
The 74 most recent commits, on their own recency sample rather than the feed above: that defaults to risk-sorted, so reusing it would plot only the top tercile and call it the spread. Big and scattered is what the model penalises. Click a dot to open it.
Repowise tracks change history across 52 files in 2233admin/cc-thinking-skills. In the last 90 days 49 files were touched, 65 times in total, most often run-calibration.js. Every commit is scored for change risk from its size, spread and the history of the files it touches.