Every commit scored for change-risk against this repo's own history, so 'elevated' means elevated here rather than on some global curve.
Needs review
20 commits sit in this repo's top risk tercile, which is 33% of the 60scored. The cut is drawn against this codebase's own history rather than a global curve, so a quiet repo still fills its top band, and here it starts at 9.6 out of 10. What pushes a commit up is size and spread together: a large change confined to one area scores below a smaller one scattered across a dozen files.
Commit categories over time, read off the subject line. Fixes carry the accent because that is the series this chart exists to show.
Began other-led, now leaning docs.
Ranked by change-risk, highest first. Priority is a tercile of this repo's own distribution, so a quiet repo still fills its top band.
| # | Commit | Author | When | Lines | Risk | Top driver |
|---|---|---|---|---|---|---|
| 1 | f6afd257bench(health-defect): add CockroachDB Repowise vs CodeScene head-to-head | RaghavChamadiya | 1mo ago | +790.9K -0 | 97%Elevated | more lines added than baseline |
| 2 | c9b8d0affeat(bench): local-model harness, defect-corpus extraction results, and runner fixes | RaghavChamadiya | 1mo ago | +835.5K -14 | 97%Elevated | more lines added than baseline |
| 3 | f01fc48bfeat(bench): same-repo tool comparison and issue-resolution-time analysis | RaghavChamadiya | 1mo ago | +681.2K -79.0K | 97%Elevated | more lines added than baseline |
| 4 | 18164c4adata(health-defect): add health scores and visualization charts | Swati Ahuja | 2mo ago | +143.8K -0 | 97%Elevated | more lines added than baseline |
| 5 | 5e0b74b2chore(harness): consolidate health-defect layout | RaghavChamadiya | 1mo ago | +7.0K -29 | 89%Elevated | more lines added than baseline |
| 6 | 8f46a9e1Agent-era defect dynamics study (28 repos, 112k commits) (#12) | Raghav Chamadiya | 1mo ago | +4.2K -0 | 89%Elevated | more lines added than baseline |
| 7 | 9f50c561data(health-defect): add canonical benchmark results for 3 repos | Swati Ahuja | 2mo ago | +10.0K -0 | 89%Elevated | more lines added than baseline |
| 8 | 295858a4Add django benchmark run and swe_qa c2/smoke test results | RaghavChamadiya | 3mo ago | +7.6K -80 | 89%Elevated | more lines added than baseline |
| 9 | 3c060ceeInitial commit: repowise-bench evaluation harness | RaghavChamadiya | 3mo ago | +33.0K -0 | 89%Elevated | more lines added than baseline |
| 10 | 63a6dfbebench(health-defect): ICSE R&R gate scripts + paper-evidence artifacts | RaghavChamadiya | 4w ago | +2.8K -0 | 84%Elevated | more lines added than baseline |
| 11 | 3d7f2ee8feat(bench): change-burst, review-coverage, and error-handling defect probes | RaghavChamadiya | 1mo ago | +1.2K -0 | 80%Elevated | more lines added than baseline |
| 12 | 3948a6c9chore(bench): track aggregate comparison outputs and clean gitignore | RaghavChamadiya | 1mo ago | +1.5K -0 | 80%Elevated | more lines added than baseline |
| 13 | c19924cdfeat(bench): statistical rigor, temporal CV, external-dataset comparison, and final report | RaghavChamadiya | 1mo ago | +1.3K -27 | 80%Elevated | more lines added than baseline |
| 14 | c13a8525feat(bench): SZZ + issue-linked defect labeling and trivial baselines | RaghavChamadiya | 2mo ago | +1.0K -124 | 80%Elevated | more lines added than baseline |
| 15 | 31c18289feat(bench): effort-aware per-hunk change-risk localization | RaghavChamadiya | 1mo ago | +853 -0 | 72%Elevated | more lines added than baseline |
| 16 | 2ba9f4ddfeat(bench): code-naturalness entropy signal and line-level localization probe | RaghavChamadiya | 1mo ago | +868 -0 | 72%Elevated | more lines added than baseline |
| 17 | 1e65b987feat(bench): candidate-feature evaluation harness + centrality probe | RaghavChamadiya | 1mo ago | +872 -0 | 72%Elevated | more lines added than baseline |
| 18 | e5aba807feat(bench): failure-forensics and size-stratified analysis tooling | RaghavChamadiya | 2mo ago | +784 -0 | 72%Elevated | more lines added than baseline |
| 19 | a8222362feat(health-defect): add core benchmark library modules | Swati Ahuja | 2mo ago | +833 -0 | 72%Elevated | more lines added than baseline |
| 20 | 592cf076Add Flask 4.8 benchmark results and clean up harness | Swati Ahuja | 3mo ago | +1.7K -3.8K | 72%Elevated | more lines added than baseline |
| 21 | a5baf114bench(flask): coherent token-reduction story — v3 (lean MCP + distill) + v2 | RaghavChamadiya | 1mo ago | +1.2K -53 | 65%Typical | more lines added than baseline |
| 22 | 94734d8bdocs(health-defect): add README and detailed benchmark report | Swati Ahuja | 2mo ago | +637 -0 | 65%Typical | more lines added than baseline |
| 23 | 51d6b9b8test(perf-detection): runtime confirmation (E4) for 7 verified findings | RaghavChamadiya | 1mo ago | +544 -6 | 61%Typical | more lines added than baseline |
| 24 | 14702d4bbench(flask): lean-MCP + long Bash/distill arms, coherent token-reduction report | RaghavChamadiya | 1mo ago | +798 -34 | 61%Typical | more lines added than baseline |
| 25 | 628dea91feat(bench): T0-anchored scoring, effort-aware metrics, and defect-calibration corpus | RaghavChamadiya | 2mo ago | +565 -467 | 61%Typical | more lines added than baseline |
| 26 | 186e37edfeat(bench): JITLine-style token model head-to-head for hunk localization | RaghavChamadiya | 1mo ago | +394 -0 | 56%Typical | more lines added than baseline |
| 27 | 075df260feat(bench): function/symbol-level SZZ defect dataset and calibration | RaghavChamadiya | 2mo ago | +470 -0 | 56%Typical | more lines added than baseline |
| 28 | 593edacdsk-learn benchmarks | Swati Ahuja | 3mo ago | +698 -3 | 56%Typical | more lines added than baseline |
| 29 | eec6d7fadocs(perf-detection): performance-risk detection technical report | RaghavChamadiya | 1mo ago | +368 -0 | 51%Typical | more lines added than baseline |
| 30 | 24dfe633bench(health-defect): expand CockroachDB head-to-head report for external readers | RaghavChamadiya | 1mo ago | +355 -112 | 51%Typical | more lines added than baseline |
| 31 | f3ec92dcfeat(bench): extend defect corpus to all nine full-tier languages | RaghavChamadiya | 1mo ago | +326 -5 | 51%Typical | more lines added than baseline |
| 32 | caf8693cdocs: rewrite the benchmark README as an accessible evidence hub | RaghavChamadiya | 4w ago | +461 -331 | 44%Typical | more lines added than baseline |
| 33 | bc5896b4feat(harness): add RefactoringMiner oracle for generated refactorings | RaghavChamadiya | 4w ago | +338 -0 | 44%Typical | more lines added than baseline |
| 34 | b69dfcadbench(health-defect): restore top-level error_analysis.py as the import root | RaghavChamadiya | 4w ago | +329 -0 | 44%Typical | more lines added than baseline |
| 35 | bef93423experiment: GAM severity shaping probe | RaghavChamadiya | 1mo ago | +341 -0 | 44%Typical | more lines added than baseline |
| 36 | 47211de1feat(health-defect): add benchmark config and orchestrator | Swati Ahuja | 2mo ago | +322 -0 | 44%Typical | more lines added than baseline |
| 37 | 868f6458bench(health-defect): G5 refit-resampling TOST + delta_boot persistence | RaghavChamadiya | 4w ago | +302 -0 | 39%Typical | more lines added than baseline |
| 38 | e363f552feat(bench): just-in-time (commit-level) defect prediction prototype | RaghavChamadiya | 1mo ago | +256 -0 | 38%Typical | more lines added than baseline |
| 39 | 7b80ca80feat(bench): calibrate the just-in-time change-risk model | RaghavChamadiya | 1mo ago | +215 -0 | 35%Typical | more lines added than baseline |
| 40 | 22c47449feat(bench): size-relative scoring experiment (research probe) | RaghavChamadiya | 2mo ago | +232 -0 | 35%Typical | more lines added than baseline |
| 41 | dfe3dfa8feat(bench): continuous coverage-gradient probe | RaghavChamadiya | 2mo ago | +188 -0 | 31%Below typical | more lines added than baseline |
| 42 | 21adfde6docs: restructure README as multi-benchmark suite overview | Swati Ahuja | 2mo ago | +193 -182 | 31%Below typical | more lines added than baseline |
| 43 | 3db7650eour own benchmarks | Swati Ahuja | 3mo ago | +358 -10 | 31%Below typical | more lines added than baseline |
| 44 | baf92671feat(bench): interpretable coverage-penalty scoring experiment | RaghavChamadiya | 2mo ago | +178 -0 | 28%Below typical | more lines added than baseline |
| 45 | 38088a53docs: copy-edit the linked benchmark reports for consistency | RaghavChamadiya | 4w ago | +173 -173 | 25%Below typical | more lines added than baseline |
| 46 | ef1d34c7docs(bench): record flask48 v2 results (stopped at 24 pairs) + raw rows | RaghavChamadiya | 1mo ago | +167 -24 | 25%Below typical | more lines added than baseline |
| 47 | d2630eebfeat(bench): wire test-coverage ingestion and report continuous-feature analysis | RaghavChamadiya | 2mo ago | +122 -8 | 23%Below typical | more lines added than baseline |
| 48 | a5791a02docs(health-defect): rewrite README for the 21-repo size-orthogonal evaluation | RaghavChamadiya | 4w ago | +117 -164 | 20%Below typical | more lines added than baseline |
| 49 | 62787a6ahealth-defect: add --nloc-cuts to error_analysis + F4 curve extractor | RaghavChamadiya | 1mo ago | +133 -1 | 20%Below typical | more lines added than baseline |
| 50 | 57e248fffeat(bench): flask48 v2 rerun — config, harness wiring, committed index, report | RaghavChamadiya | 1mo ago | +320 -8 | 17%Below typical | more lines added than baseline |
Two views of the same model: where the cuts fall, and what commit shape lands you above them.
Every scored commit, binned on the raw 0 to 10 score rather than the percentile. Percentile ranks are uniform by construction, so that axis has no shape to draw. The dashed lines are the tercile cuts behind each row's priority pill.
The 60 most recent commits, on their own recency sample rather than the feed above: that defaults to risk-sorted, so reusing it would plot only the top tercile and call it the spread. Big and scattered is what the model penalises. Click a dot to open it.
Repowise tracks change history across 120 files in repowise-dev/repowise-bench. In the last 90 days 120 files were touched, 149 times in total, most often swe_qa_runner.py. Every commit is scored for change risk from its size, spread and the history of the files it touches.