repowiserepowise
Sign in
repowise-dev/repowise-bench
OverviewDocsArchitectureKnowledge GraphFilesCode HealthRefactoring

People & History

CommitsContributorsDecisions
ChatPro
Stats
repowiserepowise
ExplorePricingDocs
Sign inIndex repoIndex your repo free
repowiserepowise-dev/repowise-bench

Commits

Every commit scored for change-risk against this repo's own history, so 'elevated' means elevated here rather than on some global curve.

Needs review

20of 60 scored

20 commits sit in this repo's top risk tercile, which is 33% of the 60scored. The cut is drawn against this codebase's own history rather than a global curve, so a quiet repo still fills its top band, and here it starts at 9.6 out of 10. What pushes a commit up is size and spread together: a large change confined to one area scores below a smaller one scattered across a dozen files.

Fix commits00%Commits whose subject reads as a bug fix rather than new work.Change diffusion1.55bitsShannon entropy of a commit's churn across its files. Zero is a single file, and every extra bit is a doubling of how widely the change spread.Review threshold9.6out of 10Score a commit has to clear to land in this repo's top tercile.

How the work changed shape

Commit categories over time, read off the subject line. Fixes carry the accent because that is the series this chart exists to show.

Began other-led, now leaning docs.

Feature47%
Refactor2%
Docs22%
Test3%
Chore5%
Other22%

Review-priority queue

Ranked by change-risk, highest first. Priority is a tercile of this repo's own distribution, so a quiet repo still fills its top band.

Commit review-priority queue
#CommitAuthorWhenLinesRiskTop driver
1
f6afd257bench(health-defect): add CockroachDB Repowise vs CodeScene head-to-head
RaghavChamadiya
1mo ago+790.9K -0
97%Elevated
more lines added than baseline
2
c9b8d0affeat(bench): local-model harness, defect-corpus extraction results, and runner fixes
RaghavChamadiya
1mo ago+835.5K -14
97%Elevated
more lines added than baseline
3
f01fc48bfeat(bench): same-repo tool comparison and issue-resolution-time analysis
RaghavChamadiya
1mo ago+681.2K -79.0K
97%Elevated
more lines added than baseline
4
18164c4adata(health-defect): add health scores and visualization charts
Swati Ahuja
2mo ago+143.8K -0
97%Elevated
more lines added than baseline
5
5e0b74b2chore(harness): consolidate health-defect layout
RaghavChamadiya
1mo ago+7.0K -29
89%Elevated
more lines added than baseline
6
8f46a9e1Agent-era defect dynamics study (28 repos, 112k commits) (#12)
Raghav Chamadiya
1mo ago+4.2K -0
89%Elevated
more lines added than baseline
7
9f50c561data(health-defect): add canonical benchmark results for 3 repos
Swati Ahuja
2mo ago+10.0K -0
89%Elevated
more lines added than baseline
8
295858a4Add django benchmark run and swe_qa c2/smoke test results
RaghavChamadiya
3mo ago+7.6K -80
89%Elevated
more lines added than baseline
9
3c060ceeInitial commit: repowise-bench evaluation harness
RaghavChamadiya
3mo ago+33.0K -0
89%Elevated
more lines added than baseline
10
63a6dfbebench(health-defect): ICSE R&R gate scripts + paper-evidence artifacts
RaghavChamadiya
4w ago+2.8K -0
84%Elevated
more lines added than baseline
11
3d7f2ee8feat(bench): change-burst, review-coverage, and error-handling defect probes
RaghavChamadiya
1mo ago+1.2K -0
80%Elevated
more lines added than baseline
12
3948a6c9chore(bench): track aggregate comparison outputs and clean gitignore
RaghavChamadiya
1mo ago+1.5K -0
80%Elevated
more lines added than baseline
13
c19924cdfeat(bench): statistical rigor, temporal CV, external-dataset comparison, and final report
RaghavChamadiya
1mo ago+1.3K -27
80%Elevated
more lines added than baseline
14
c13a8525feat(bench): SZZ + issue-linked defect labeling and trivial baselines
RaghavChamadiya
2mo ago+1.0K -124
80%Elevated
more lines added than baseline
15
31c18289feat(bench): effort-aware per-hunk change-risk localization
RaghavChamadiya
1mo ago+853 -0
72%Elevated
more lines added than baseline
16
2ba9f4ddfeat(bench): code-naturalness entropy signal and line-level localization probe
RaghavChamadiya
1mo ago+868 -0
72%Elevated
more lines added than baseline
17
1e65b987feat(bench): candidate-feature evaluation harness + centrality probe
RaghavChamadiya
1mo ago+872 -0
72%Elevated
more lines added than baseline
18
e5aba807feat(bench): failure-forensics and size-stratified analysis tooling
RaghavChamadiya
1mo ago+784 -0
72%Elevated
more lines added than baseline
19
a8222362feat(health-defect): add core benchmark library modules
Swati Ahuja
2mo ago+833 -0
72%Elevated
more lines added than baseline
20
592cf076Add Flask 4.8 benchmark results and clean up harness
Swati Ahuja
3mo ago+1.7K -3.8K
72%Elevated
more lines added than baseline
21
a5baf114bench(flask): coherent token-reduction story — v3 (lean MCP + distill) + v2
RaghavChamadiya
1mo ago+1.2K -53
65%Typical
more lines added than baseline
22
94734d8bdocs(health-defect): add README and detailed benchmark report
Swati Ahuja
2mo ago+637 -0
65%Typical
more lines added than baseline
23
51d6b9b8test(perf-detection): runtime confirmation (E4) for 7 verified findings
RaghavChamadiya
1mo ago+544 -6
61%Typical
more lines added than baseline
24
14702d4bbench(flask): lean-MCP + long Bash/distill arms, coherent token-reduction report
RaghavChamadiya
1mo ago+798 -34
61%Typical
more lines added than baseline
25
628dea91feat(bench): T0-anchored scoring, effort-aware metrics, and defect-calibration corpus
RaghavChamadiya
2mo ago+565 -467
61%Typical
more lines added than baseline
26
186e37edfeat(bench): JITLine-style token model head-to-head for hunk localization
RaghavChamadiya
1mo ago+394 -0
56%Typical
more lines added than baseline
27
075df260feat(bench): function/symbol-level SZZ defect dataset and calibration
RaghavChamadiya
2mo ago+470 -0
56%Typical
more lines added than baseline
28
593edacdsk-learn benchmarks
Swati Ahuja
3mo ago+698 -3
56%Typical
more lines added than baseline
29
eec6d7fadocs(perf-detection): performance-risk detection technical report
RaghavChamadiya
1mo ago+368 -0
51%Typical
more lines added than baseline
30
24dfe633bench(health-defect): expand CockroachDB head-to-head report for external readers
RaghavChamadiya
1mo ago+355 -112
51%Typical
more lines added than baseline
31
f3ec92dcfeat(bench): extend defect corpus to all nine full-tier languages
RaghavChamadiya
1mo ago+326 -5
51%Typical
more lines added than baseline
32
caf8693cdocs: rewrite the benchmark README as an accessible evidence hub
RaghavChamadiya
4w ago+461 -331
44%Typical
more lines added than baseline
33
bc5896b4feat(harness): add RefactoringMiner oracle for generated refactorings
RaghavChamadiya
4w ago+338 -0
44%Typical
more lines added than baseline
34
b69dfcadbench(health-defect): restore top-level error_analysis.py as the import root
RaghavChamadiya
4w ago+329 -0
44%Typical
more lines added than baseline
35
bef93423experiment: GAM severity shaping probe
RaghavChamadiya
1mo ago+341 -0
44%Typical
more lines added than baseline
36
47211de1feat(health-defect): add benchmark config and orchestrator
Swati Ahuja
2mo ago+322 -0
44%Typical
more lines added than baseline
37
868f6458bench(health-defect): G5 refit-resampling TOST + delta_boot persistence
RaghavChamadiya
4w ago+302 -0
39%Typical
more lines added than baseline
38
e363f552feat(bench): just-in-time (commit-level) defect prediction prototype
RaghavChamadiya
1mo ago+256 -0
38%Typical
more lines added than baseline
39
7b80ca80feat(bench): calibrate the just-in-time change-risk model
RaghavChamadiya
1mo ago+215 -0
35%Typical
more lines added than baseline
40
22c47449feat(bench): size-relative scoring experiment (research probe)
RaghavChamadiya
1mo ago+232 -0
35%Typical
more lines added than baseline
41
dfe3dfa8feat(bench): continuous coverage-gradient probe
RaghavChamadiya
1mo ago+188 -0
31%Below typical
more lines added than baseline
42
21adfde6docs: restructure README as multi-benchmark suite overview
Swati Ahuja
2mo ago+193 -182
31%Below typical
more lines added than baseline
43
3db7650eour own benchmarks
Swati Ahuja
3mo ago+358 -10
31%Below typical
more lines added than baseline
44
baf92671feat(bench): interpretable coverage-penalty scoring experiment
RaghavChamadiya
1mo ago+178 -0
28%Below typical
more lines added than baseline
45
38088a53docs: copy-edit the linked benchmark reports for consistency
RaghavChamadiya
4w ago+173 -173
25%Below typical
more lines added than baseline
46
ef1d34c7docs(bench): record flask48 v2 results (stopped at 24 pairs) + raw rows
RaghavChamadiya
1mo ago+167 -24
25%Below typical
more lines added than baseline
47
d2630eebfeat(bench): wire test-coverage ingestion and report continuous-feature analysis
RaghavChamadiya
2mo ago+122 -8
23%Below typical
more lines added than baseline
48
a5791a02docs(health-defect): rewrite README for the 21-repo size-orthogonal evaluation
RaghavChamadiya
4w ago+117 -164
20%Below typical
more lines added than baseline
49
62787a6ahealth-defect: add --nloc-cuts to error_analysis + F4 curve extractor
RaghavChamadiya
1mo ago+133 -1
20%Below typical
more lines added than baseline
50
57e248fffeat(bench): flask48 v2 rerun — config, harness wiring, committed index, report
RaghavChamadiya
1mo ago+320 -8
17%Below typical
more lines added than baseline
  • f6afd257bench(health-defect): add CockroachDB Repowise vs CodeScene head-to-head
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +790.9K -0
    Risk
    97%Elevated
  • c9b8d0affeat(bench): local-model harness, defect-corpus extraction results, and runner fixes
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +835.5K -14
    Risk
    97%Elevated
  • f01fc48bfeat(bench): same-repo tool comparison and issue-resolution-time analysis
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +681.2K -79.0K
    Risk
    97%Elevated
  • 18164c4adata(health-defect): add health scores and visualization charts
    Author
    Swati Ahuja
    When
    2mo ago
    Lines
    +143.8K -0
    Risk
    97%Elevated
  • 5e0b74b2chore(harness): consolidate health-defect layout
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +7.0K -29
    Risk
    89%Elevated
  • 8f46a9e1Agent-era defect dynamics study (28 repos, 112k commits) (#12)
    Author
    Raghav Chamadiya
    When
    1mo ago
    Lines
    +4.2K -0
    Risk
    89%Elevated
  • 9f50c561data(health-defect): add canonical benchmark results for 3 repos
    Author
    Swati Ahuja
    When
    2mo ago
    Lines
    +10.0K -0
    Risk
    89%Elevated
  • 295858a4Add django benchmark run and swe_qa c2/smoke test results
    Author
    RaghavChamadiya
    When
    3mo ago
    Lines
    +7.6K -80
    Risk
    89%Elevated
  • 3c060ceeInitial commit: repowise-bench evaluation harness
    Author
    RaghavChamadiya
    When
    3mo ago
    Lines
    +33.0K -0
    Risk
    89%Elevated
  • 63a6dfbebench(health-defect): ICSE R&R gate scripts + paper-evidence artifacts
    Author
    RaghavChamadiya
    When
    4w ago
    Lines
    +2.8K -0
    Risk
    84%Elevated
  • 3d7f2ee8feat(bench): change-burst, review-coverage, and error-handling defect probes
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +1.2K -0
    Risk
    80%Elevated
  • 3948a6c9chore(bench): track aggregate comparison outputs and clean gitignore
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +1.5K -0
    Risk
    80%Elevated
  • c19924cdfeat(bench): statistical rigor, temporal CV, external-dataset comparison, and final report
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +1.3K -27
    Risk
    80%Elevated
  • c13a8525feat(bench): SZZ + issue-linked defect labeling and trivial baselines
    Author
    RaghavChamadiya
    When
    2mo ago
    Lines
    +1.0K -124
    Risk
    80%Elevated
  • 31c18289feat(bench): effort-aware per-hunk change-risk localization
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +853 -0
    Risk
    72%Elevated
  • 2ba9f4ddfeat(bench): code-naturalness entropy signal and line-level localization probe
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +868 -0
    Risk
    72%Elevated
  • 1e65b987feat(bench): candidate-feature evaluation harness + centrality probe
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +872 -0
    Risk
    72%Elevated
  • e5aba807feat(bench): failure-forensics and size-stratified analysis tooling
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +784 -0
    Risk
    72%Elevated
  • a8222362feat(health-defect): add core benchmark library modules
    Author
    Swati Ahuja
    When
    2mo ago
    Lines
    +833 -0
    Risk
    72%Elevated
  • 592cf076Add Flask 4.8 benchmark results and clean up harness
    Author
    Swati Ahuja
    When
    3mo ago
    Lines
    +1.7K -3.8K
    Risk
    72%Elevated
  • a5baf114bench(flask): coherent token-reduction story — v3 (lean MCP + distill) + v2
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +1.2K -53
    Risk
    65%Typical
  • 94734d8bdocs(health-defect): add README and detailed benchmark report
    Author
    Swati Ahuja
    When
    2mo ago
    Lines
    +637 -0
    Risk
    65%Typical
  • 51d6b9b8test(perf-detection): runtime confirmation (E4) for 7 verified findings
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +544 -6
    Risk
    61%Typical
  • 14702d4bbench(flask): lean-MCP + long Bash/distill arms, coherent token-reduction report
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +798 -34
    Risk
    61%Typical
  • 628dea91feat(bench): T0-anchored scoring, effort-aware metrics, and defect-calibration corpus
    Author
    RaghavChamadiya
    When
    2mo ago
    Lines
    +565 -467
    Risk
    61%Typical
  • 186e37edfeat(bench): JITLine-style token model head-to-head for hunk localization
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +394 -0
    Risk
    56%Typical
  • 075df260feat(bench): function/symbol-level SZZ defect dataset and calibration
    Author
    RaghavChamadiya
    When
    2mo ago
    Lines
    +470 -0
    Risk
    56%Typical
  • 593edacdsk-learn benchmarks
    Author
    Swati Ahuja
    When
    3mo ago
    Lines
    +698 -3
    Risk
    56%Typical
  • eec6d7fadocs(perf-detection): performance-risk detection technical report
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +368 -0
    Risk
    51%Typical
  • 24dfe633bench(health-defect): expand CockroachDB head-to-head report for external readers
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +355 -112
    Risk
    51%Typical
  • f3ec92dcfeat(bench): extend defect corpus to all nine full-tier languages
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +326 -5
    Risk
    51%Typical
  • caf8693cdocs: rewrite the benchmark README as an accessible evidence hub
    Author
    RaghavChamadiya
    When
    4w ago
    Lines
    +461 -331
    Risk
    44%Typical
  • bc5896b4feat(harness): add RefactoringMiner oracle for generated refactorings
    Author
    RaghavChamadiya
    When
    4w ago
    Lines
    +338 -0
    Risk
    44%Typical
  • b69dfcadbench(health-defect): restore top-level error_analysis.py as the import root
    Author
    RaghavChamadiya
    When
    4w ago
    Lines
    +329 -0
    Risk
    44%Typical
  • bef93423experiment: GAM severity shaping probe
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +341 -0
    Risk
    44%Typical
  • 47211de1feat(health-defect): add benchmark config and orchestrator
    Author
    Swati Ahuja
    When
    2mo ago
    Lines
    +322 -0
    Risk
    44%Typical
  • 868f6458bench(health-defect): G5 refit-resampling TOST + delta_boot persistence
    Author
    RaghavChamadiya
    When
    4w ago
    Lines
    +302 -0
    Risk
    39%Typical
  • e363f552feat(bench): just-in-time (commit-level) defect prediction prototype
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +256 -0
    Risk
    38%Typical
  • 7b80ca80feat(bench): calibrate the just-in-time change-risk model
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +215 -0
    Risk
    35%Typical
  • 22c47449feat(bench): size-relative scoring experiment (research probe)
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +232 -0
    Risk
    35%Typical
  • dfe3dfa8feat(bench): continuous coverage-gradient probe
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +188 -0
    Risk
    31%Below typical
  • 21adfde6docs: restructure README as multi-benchmark suite overview
    Author
    Swati Ahuja
    When
    2mo ago
    Lines
    +193 -182
    Risk
    31%Below typical
  • 3db7650eour own benchmarks
    Author
    Swati Ahuja
    When
    3mo ago
    Lines
    +358 -10
    Risk
    31%Below typical
  • baf92671feat(bench): interpretable coverage-penalty scoring experiment
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +178 -0
    Risk
    28%Below typical
  • 38088a53docs: copy-edit the linked benchmark reports for consistency
    Author
    RaghavChamadiya
    When
    4w ago
    Lines
    +173 -173
    Risk
    25%Below typical
  • ef1d34c7docs(bench): record flask48 v2 results (stopped at 24 pairs) + raw rows
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +167 -24
    Risk
    25%Below typical
  • d2630eebfeat(bench): wire test-coverage ingestion and report continuous-feature analysis
    Author
    RaghavChamadiya
    When
    2mo ago
    Lines
    +122 -8
    Risk
    23%Below typical
  • a5791a02docs(health-defect): rewrite README for the 21-repo size-orthogonal evaluation
    Author
    RaghavChamadiya
    When
    4w ago
    Lines
    +117 -164
    Risk
    20%Below typical
  • 62787a6ahealth-defect: add --nloc-cuts to error_analysis + F4 curve extractor
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +133 -1
    Risk
    20%Below typical
  • 57e248fffeat(bench): flask48 v2 rerun — config, harness wiring, committed index, report
    Author
    RaghavChamadiya
    When
    1mo ago
    Lines
    +320 -8
    Risk
    17%Below typical
Showing 50 of 60 commits

How the score behaves here

Change risk →

Two views of the same model: where the cuts fall, and what commit shape lands you above them.

Score distribution

Every scored commit, binned on the raw 0 to 10 score rather than the percentile. Percentile ranks are uniform by construction, so that axis has no shape to draw. The dashed lines are the tercile cuts behind each row's priority pill.

022↑ typical↑ elevated0.05.010.0Change-risk score →
Below typical
Typical
Elevated

Size against diffusion

The 60 most recent commits, on their own recency sample rather than the feed above: that defaults to risk-sorted, so reusing it would plot only the top tercile and call it the spread. Big and scattered is what the model penalises. Click a dot to open it.

1101001,00010,000100,000Lines changed (log) →0.06.3Diffusion →
Below typical20
Typical20
Elevated20

Commit history for repowise-dev/repowise-bench

Repowise tracks change history across 120 files in repowise-dev/repowise-bench. In the last 90 days 120 files were touched, 149 times in total, most often swe_qa_runner.py. Every commit is scored for change risk from its size, spread and the history of the files it touches.