Bump the github-actions-dependencies group across 1 directory with 2 updates - #963
Conversation
📊 Skill Evaluation Results10 skill(s) evaluated — 6 improved, 4 no credible improvement. A skill passes only on a credible improvement over baseline (mean preference > 0 with its 95% CI above 0);
ℹ️ Column legend
✅ create-datadriven-aspnetcore — detailsReason: Mean preference +56.0% [95% CI 2.2%, 109.8%], win rate 80.0% (4W/1T/0L over 5 trial(s)) — credibly better
✅ dotnet-maui-doctor — detailsReason: Mean preference +70.0% [95% CI 43.2%, 96.8%], win rate 100.0% (8W/0T/0L over 8 trial(s)) — credibly better
✅ maui-app-lifecycle — detailsReason: Mean preference +40.0% [95% CI 40.0%, 40.0%], win rate 100.0% (4W/0T/0L over 4 trial(s)) — credibly better
❌ maui-collectionview — detailsReason: Mean preference +30.0% [95% CI -1.8%, 61.8%], win rate 75.0% (3W/1T/0L over 4 trial(s)) — not credible (95% CI includes 0)
✅ maui-data-binding — detailsReason: Mean preference +40.0% [95% CI 40.0%, 40.0%], win rate 100.0% (4W/0T/0L over 4 trial(s)) — credibly better
❌ maui-dependency-injection — detailsReason: Mean preference +20.0% [95% CI -43.6%, 83.6%], win rate 75.0% (3W/0T/1L over 4 trial(s)) — not credible (95% CI includes 0)
✅ maui-safe-area — detailsReason: Mean preference +100.0% [95% CI 100.0%, 100.0%], win rate 100.0% (4W/0T/0L over 4 trial(s)) — credibly better
✅ maui-shell-navigation — detailsReason: Mean preference +55.0% [95% CI 7.3%, 102.7%], win rate 100.0% (4W/0T/0L over 4 trial(s)) — credibly better
❌ maui-theming — detailsReason: Mean preference +10.0% [95% CI -21.8%, 41.8%], win rate 25.0% (1W/3T/0L over 4 trial(s)) — not credible (95% CI includes 0)
❌ optimizing-ef-core-queries — detailsReason: Mean preference -40.0% [95% CI -40.0%, -40.0%], win rate 0.0% (0W/0T/1L over 1 trial(s)) — no improvement
🔍 Full Results - additional metrics and failure investigation steps
▶ Sessions Visualisation -- interactive replay of all evaluation sessions |
|
✅ Evaluation passed for |
b6213a8 to
63cfc57
Compare
|
❌ Evaluation ran but produced no results. The evaluate job completed but no |
|
✅ Evaluation passed for |
63cfc57 to
fe76869
Compare
📊 Skill Evaluation Results27 skill(s) evaluated — ✅ 2 improved, ❌ 0 no credible change, 🔻 0 regressed.
A skill passes only on a credible net win over baseline: more wins than losses, by an exact one-sided sign test at
ℹ️ Column legend
|
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Diagnose build failures from binlog only (no source files) | +0.0% | +0.0% | 0/1/0 |
⚠️ binlog-generation — details
Reason: Net win +100.0% (3W/0T/0L over 3 trial(s), sign test p=0.125), mean preference +80.0% — underpowered (3 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Build multiple configurations with unique binlogs | +100.0% | +100.0% | 1/0/0 |
| ▲ Build project with /bl flag | +100.0% | +100.0% | 1/0/0 |
| ▲ Build with /bl in PowerShell | +100.0% | +40.0% | 1/0/0 |
⚠️ build-parallelism — details
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Analyze build parallelism bottlenecks | +100.0% | +40.0% | 1/0/0 |
⚠️ build-perf-baseline — details
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Establish build performance baseline and recommend optimizations | +100.0% | +40.0% | 1/0/0 |
⚠️ build-perf-diagnostics — details
Reason: Net win +0.0% (0W/1T/0L over 1 trial(s), sign test p=1.000), mean preference +0.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Diagnose slow build for a small project | +0.0% | +0.0% | 0/1/0 |
⚠️ check-bin-obj-clash — details
Reason: Net win +0.0% (0W/1T/0L over 1 trial(s), sign test p=1.000), mean preference +0.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Diagnose bin/obj output path clashes | +0.0% | +0.0% | 0/1/0 |
⚠️ collect-user-input — details
Reason: Net win +50.0% (1W/1T/0L over 2 trial(s), sign test p=0.500), mean preference +20.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Event registration with custom validation | +100.0% | +40.0% | 1/0/0 |
| = Multi-step booking form with cross-field validation | +0.0% | +0.0% | 0/1/0 |
⚠️ configure-auth — details
Reason: Net win +50.0% (1W/1T/0L over 2 trial(s), sign test p=0.500), mean preference +20.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Login and account management in a globally interactive app | +0.0% | +0.0% | 0/1/0 |
| ▲ Multi-tier app with WebAssembly auth | +100.0% | +40.0% | 1/0/0 |
⚠️ coordinate-components — details
Reason: Net win +50.0% (1W/1T/0L over 2 trial(s), sign test p=0.500), mean preference +50.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Multi-tenant notification hub with cross-component fan-out | +100.0% | +100.0% | 1/0/0 |
| = Warehouse dashboard with site selector and live stock alerts | +0.0% | +0.0% | 0/1/0 |
⚠️ create-blazor-project — details
Reason: Net win +100.0% (3W/0T/0L over 3 trial(s), sign test p=0.125), mean preference +60.0% — underpowered (3 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Global logistics tracking for worldwide users | +100.0% | +100.0% | 1/0/0 |
| ▲ Recipe community with interactive ratings on static pages | +100.0% | +40.0% | 1/0/0 |
| ▲ University course catalog with enrollment form | +100.0% | +40.0% | 1/0/0 |
⚠️ directory-build-organization — details
Reason: Net win +0.0% (0W/1T/0L over 1 trial(s), sign test p=1.000), mean preference +0.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Organize build infrastructure for a multi-project repo | +0.0% | +0.0% | 0/1/0 |
⚠️ eval-performance — details
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Analyze MSBuild evaluation performance issues | +100.0% | +40.0% | 1/0/0 |
⚠️ extension-points — details
Reason: Net win +100.0% (3W/0T/0L over 3 trial(s), sign test p=0.125), mean preference +40.0% — underpowered (3 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Diagnose NuGet package and repo extension conflicts | +100.0% | +40.0% | 1/0/0 |
| ▲ Diagnose build extension point failures | +100.0% | +40.0% | 1/0/0 |
| ▲ Fix extension point anti-patterns | +100.0% | +40.0% | 1/0/0 |
⚠️ fetch-and-send-data — details
Reason: Net win +100.0% (2W/0T/0L over 2 trial(s), sign test p=0.250), mean preference +40.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Real-time shipment tracker with Auto interactivity | +100.0% | +40.0% | 1/0/0 |
| ▲ Recipe browser with resilient data loading | +100.0% | +40.0% | 1/0/0 |
⚠️ including-generated-files — details
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Diagnose generated file inclusion failure | +100.0% | +40.0% | 1/0/0 |
⚠️ incremental-build — details
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Analyze incremental build issues | +100.0% | +40.0% | 1/0/0 |
⚠️ item-management — details
Reason: Net win +0.0% (1W/1T/1L over 3 trial(s), sign test p=0.750), mean preference +0.0% — underpowered (3 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▼ Diagnose cascading item and batching bugs in code generation pipeline | -100.0% | -40.0% | 0/0/1 |
| = Diagnose item group and batching issues | +0.0% | +0.0% | 0/1/0 |
| ▲ Fix item management anti-patterns | +100.0% | +40.0% | 1/0/0 |
⚠️ msbuild-antipatterns — details
Reason: Net win +0.0% (0W/4T/0L over 4 trial(s), sign test p=1.000), mean preference +0.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Add a module to an F# project | +0.0% | +0.0% | 0/1/0 |
| = Add a signature file to define public API | +0.0% | +0.0% | 0/1/0 |
| = Fix broken file order causing FS0039 | +0.0% | +0.0% | 0/1/0 |
| = Review MSBuild files for anti-patterns and style issues | +0.0% | +0.0% | 0/1/0 |
⚠️ msbuild-modernization — details
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Modernize legacy project to SDK-style | +100.0% | +40.0% | 1/0/0 |
⚠️ msbuild-server — details
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Recommend MSBuild Server for slow CLI incremental builds | +100.0% | +40.0% | 1/0/0 |
⚠️ property-patterns — details
Reason: Net win -33.3% (0W/2T/1L over 3 trial(s), sign test p=0.500), mean preference -13.3% — underpowered (3 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Diagnose multi-level property hierarchy bugs | +0.0% | +0.0% | 0/1/0 |
| = Diagnose shared build property issues | +0.0% | +0.0% | 0/1/0 |
| ▼ Fix shared property configuration | -100.0% | -40.0% | 0/0/1 |
⚠️ resolve-project-references — details
Reason: Net win +0.0% (0W/1T/0L over 1 trial(s), sign test p=1.000), mean preference +0.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Explain misleading ResolveProjectReferences time | +0.0% | +0.0% | 0/1/0 |
⚠️ support-prerendering — details
Reason: Net win +100.0% (2W/0T/0L over 2 trial(s), sign test p=0.250), mean preference +70.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Equipment inventory loaded once | +100.0% | +100.0% | 1/0/0 |
| ▲ Notifications page with live polling | +100.0% | +40.0% | 1/0/0 |
⚠️ target-authoring — details
Reason: Net win +100.0% (3W/0T/0L over 3 trial(s), sign test p=0.125), mean preference +40.0% — underpowered (3 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Diagnose broken SDK target chain across files | +100.0% | +40.0% | 1/0/0 |
| ▲ Diagnose custom target build regression | +100.0% | +40.0% | 1/0/0 |
| ▲ Fix custom target anti-patterns | +100.0% | +40.0% | 1/0/0 |
⚠️ use-js-interop — details
Reason: Net win +100.0% (4W/0T/0L over 4 trial(s), sign test p=0.063), mean preference +40.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Auto-saving notepad that survives page reloads | +100.0% | +40.0% | 1/0/0 |
| ▲ Infinite scroll list using IntersectionObserver | +100.0% | +40.0% | 1/0/0 |
| ▲ Responsive layout that adapts to screen size | +100.0% | +40.0% | 1/0/0 |
| ▲ User activity tracker that detects idle timeout | +100.0% | +40.0% | 1/0/0 |
Per-scenario details for 2 skill(s) were omitted to keep this comment under GitHub's 65,536-character limit — open the job's step summary or Full Results for the complete breakdown.
🔍 Full Results - additional metrics and failure investigation steps
▶ Sessions Visualisation -- interactive replay of all evaluation sessions
📊 Session Analytics (preview) -- aggregated metrics across evaluation sessions
…updates Bumps the github-actions-dependencies group with 2 updates in the / directory: [actions/checkout](https://github.com/actions/checkout) and [actions/setup-python](https://github.com/actions/setup-python). Updates `actions/checkout` from 7.0.0 to 7.0.1 - [Release notes](https://github.com/actions/checkout/releases) - [Changelog](https://github.com/actions/checkout/blob/main/CHANGELOG.md) - [Commits](actions/checkout@v7...3d3c42e) Updates `actions/setup-python` from 5.6.0 to 7.0.0 - [Release notes](https://github.com/actions/setup-python/releases) - [Commits](actions/setup-python@a26af69...5fda3b9) --- updated-dependencies: - dependency-name: actions/checkout dependency-version: 7.0.1 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: github-actions-dependencies - dependency-name: actions/setup-python dependency-version: 7.0.0 dependency-type: direct:production update-type: version-update:semver-major dependency-group: github-actions-dependencies ... Signed-off-by: dependabot[bot] <support@github.com>
fe76869 to
adbbc11
Compare
📊 Skill Evaluation Results36 skill(s) evaluated — ✅ 0 improved, ❌ 0 no credible change, 🔻 0 regressed.
A skill passes only on a credible net win over baseline: more wins than losses, by an exact one-sided sign test at
ℹ️ Column legend
|
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Distinguish weak and meaningful assertions in a pytest suite | +0.0% | +0.0% | 0/0/0 |
| = Flag assertion-free tests and trivial-only assertions | +0.0% | +0.0% | 0/0/0 |
| = Identify low assertion diversity in equality-dominated test suite | +0.0% | +0.0% | 0/0/0 |
| = Identify self-referential assertions in identity and round-trip tests | +0.0% | +0.0% | 0/0/0 |
| = Judge assertion strength in a shallow Jest suite | +0.0% | +0.0% | 0/0/0 |
| = Recognize well-diversified assertion usage | +0.0% | +0.0% | 0/0/0 |
⚠️ binlog-failure-analysis — details
Reason: Net win +0.0% (0W/1T/0L over 1 trial(s), sign test p=1.000), mean preference +0.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Diagnose build failures from binlog only (no source files) | +0.0% | +0.0% | 0/1/0 |
⚠️ binlog-generation — details
Reason: Net win +100.0% (3W/0T/0L over 3 trial(s), sign test p=0.125), mean preference +100.0% — underpowered (3 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Build multiple configurations with unique binlogs | +100.0% | +100.0% | 1/0/0 |
| ▲ Build project with /bl flag | +100.0% | +100.0% | 1/0/0 |
| ▲ Build with /bl in PowerShell | +100.0% | +100.0% | 1/0/0 |
⚠️ build-parallelism — details
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Analyze build parallelism bottlenecks | +100.0% | +40.0% | 1/0/0 |
⚠️ build-perf-baseline — details
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Establish build performance baseline and recommend optimizations | +100.0% | +40.0% | 1/0/0 |
⚠️ build-perf-diagnostics — details
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Diagnose slow build for a small project | +100.0% | +40.0% | 1/0/0 |
⚠️ check-bin-obj-clash — details
Reason: Net win -100.0% (0W/0T/1L over 1 trial(s), sign test p=0.500), mean preference -40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▼ Diagnose bin/obj output path clashes | -100.0% | -40.0% | 0/0/1 |
⚠️ code-testing-agent — details
Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 10 errored — inconclusive (comparison errors)
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Diagnose failing tests without generating a new suite | +0.0% | +0.0% | 0/0/0 |
| = Does not revert a gutted-looking workspace (workspace integrity) | +0.0% | +0.0% | 0/0/0 |
| = Extend an existing suite to the untested method only | +0.0% | +0.0% | 0/0/0 |
| = Generate Vitest tests for the shopping-cart library (TypeScript polyglot) | +0.0% | +0.0% | 0/0/0 |
| = Keep a single-function request proportional | +0.0% | +0.0% | 0/0/0 |
⚠️ coverage-analysis — details
Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 21 errored — inconclusive (comparison errors)
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Account for a gap spread across several members | +0.0% | +0.0% | 0/0/0 |
| = Analyse a CI Cobertura report without re-running tests or installing tools | +0.0% | +0.0% | 0/0/0 |
| = Coverage plateau diagnosis | +0.0% | +0.0% | 0/0/0 |
| = Distinguish branch coverage from line coverage | +0.0% | +0.0% | 0/0/0 |
| = Project-wide coverage analysis with existing Cobertura data | +0.0% | +0.0% | 0/0/0 |
| = Refactoring safety assessment from coverage data | +0.0% | +0.0% | 0/0/0 |
| = Run coverage from scratch without existing data | +0.0% | +0.0% | 0/0/0 |
⚠️ crap-score — details
Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 5 errored, 1 unmatched — inconclusive (comparison errors)
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Calculate CRAP score for a single method with partial coverage | +0.0% | +0.0% | 0/0/0 |
| = Generate coverage then compute CRAP score | +0.0% | +0.0% | 0/0/0 |
| = Identify riskiest methods across a file | +0.0% | +0.0% | 0/0/0 |
| = Recognize when complexity alone blocks the CRAP threshold | +0.0% | +0.0% | 0/0/0 |
| = Recompute complexity instead of trusting a stale source comment | +0.0% | +0.0% | 0/0/0 |
| = Report a fully covered method at its complexity floor | +0.0% | +0.0% | 0/0/0 |
⚠️ detect-static-dependencies — details
Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 7 errored — inconclusive (comparison errors)
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Decline scan for non-C# project | +0.0% | +0.0% | 0/0/0 |
| = Detect statics inside lambda expressions and LINQ queries | +0.0% | +0.0% | 0/0/0 |
| = Detect time-related statics and recommend TimeProvider | +0.0% | +0.0% | 0/0/0 |
| = Exclude obj and bin directories from the scan | +0.0% | +0.0% | 0/0/0 |
| = Identify static dependencies in a multi-class project | +0.0% | +0.0% | 0/0/0 |
| = Keep one authoritative total with file line locations and no seam for pure helpers | +0.0% | +0.0% | 0/0/0 |
| = Verify structured report includes file count, categories, and top patterns | +0.0% | +0.0% | 0/0/0 |
⚠️ directory-build-organization — details
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Organize build infrastructure for a multi-project repo | +100.0% | +40.0% | 1/0/0 |
⚠️ eval-performance — details
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Analyze MSBuild evaluation performance issues | +100.0% | +40.0% | 1/0/0 |
⚠️ extension-points — details
Reason: Net win +33.3% (2W/0T/1L over 3 trial(s), sign test p=0.500), mean preference +13.3% — underpowered (3 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Diagnose NuGet package and repo extension conflicts | +100.0% | +40.0% | 1/0/0 |
| ▲ Diagnose build extension point failures | +100.0% | +40.0% | 1/0/0 |
| ▼ Fix extension point anti-patterns | -100.0% | -40.0% | 0/0/1 |
⚠️ filter-syntax — details
Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 5 errored — inconclusive (comparison errors)
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Filter a TUnit suite down to one class and one property value | +0.0% | +0.0% | 0/0/0 |
| = Filter xUnit v3 tests that do not accept the generic filter expression | +0.0% | +0.0% | 0/0/0 |
| = Pass a filter to a Microsoft.Testing.Platform project on the .NET 9 SDK | +0.0% | +0.0% | 0/0/0 |
| = Select one category and exclude another on a VSTest project | +0.0% | +0.0% | 0/0/0 |
| = Translate CI filter expressions after moving to xUnit v3 | +0.0% | +0.0% | 0/0/0 |
⚠️ find-untested-sources — details
Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 18 errored — inconclusive (comparison errors)
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Choose the polyglot engine for a mixed C# and TypeScript repository | +0.0% | +0.0% | 0/0/0 |
| = Disambiguate duplicate C# type names by namespace | +0.0% | +0.0% | 0/0/0 |
| = Disambiguate duplicate C# types by nested test namespace | +0.0% | +0.0% | 0/0/0 |
| = Exclude generated sources and surface an orphan test | +0.0% | +0.0% | 0/0/0 |
| = Identify an unpaired TypeScript module | +0.0% | +0.0% | 0/0/0 |
| = Pair sources to tests across a src/tests directory split | +0.0% | +0.0% | 0/0/0 |
⚠️ generate-testability-wrappers — details
Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 15 errored — inconclusive (comparison errors)
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Decline wrapper generation for already-abstracted code | +0.0% | +0.0% | 0/0/0 |
| = Generate TimeProvider adoption for DateTime.UtcNow | +0.0% | +0.0% | 0/0/0 |
| = Generate custom Environment wrapper | +0.0% | +0.0% | 0/0/0 |
| = Make time controllable in a library that has no DI container | +0.0% | +0.0% | 0/0/0 |
| = Recommend System.IO.Abstractions for file system calls | +0.0% | +0.0% | 0/0/0 |
⚠️ grade-tests — details
Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 18 errored — inconclusive (comparison errors)
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Ask for a bounded list instead of grading the workspace | +0.0% | +0.0% | 0/0/0 |
| = Grade C# tests against available production code | +0.0% | +0.0% | 0/0/0 |
| = Grade Go table-driven tests without misreading the loop as branching | +0.0% | +0.0% | 0/0/0 |
| = Grade pytest test methods using the same rubric | +0.0% | +0.0% | 0/0/0 |
| = Grade tests when the production code under test is unavailable | +0.0% | +0.0% | 0/0/0 |
| = Keep a 62-test grading report readable as a PR comment | +0.0% | +0.0% | 0/0/0 |
⚠️ including-generated-files — details
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Diagnose generated file inclusion failure | +100.0% | +40.0% | 1/0/0 |
⚠️ incremental-build — details
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Analyze incremental build issues | +100.0% | +40.0% | 1/0/0 |
⚠️ item-management — details
Reason: Net win +100.0% (3W/0T/0L over 3 trial(s), sign test p=0.125), mean preference +40.0% — underpowered (3 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Diagnose cascading item and batching bugs in code generation pipeline | +100.0% | +40.0% | 1/0/0 |
| ▲ Diagnose item group and batching issues | +100.0% | +40.0% | 1/0/0 |
| ▲ Fix item management anti-patterns | +100.0% | +40.0% | 1/0/0 |
⚠️ migrate-static-to-wrapper — details
Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 6 errored — inconclusive (comparison errors)
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Add the required using directive and update tests with a test double | +0.0% | +0.0% | 0/0/0 |
| = Decline migration when wrapper does not exist yet | +0.0% | +0.0% | 0/0/0 |
| = Migrate DateTime.UtcNow to TimeProvider in a service class | +0.0% | +0.0% | 0/0/0 |
| = Migrate a static helper class without breaking its callers | +0.0% | +0.0% | 0/0/0 |
| = Migrate only in scoped files, leaving others untouched | +0.0% | +0.0% | 0/0/0 |
| = Preserve DateTimeKind when migrating to TimeProvider | +0.0% | +0.0% | 0/0/0 |
⚠️ msbuild-antipatterns — details
Reason: Net win +25.0% (1W/3T/0L over 4 trial(s), sign test p=0.500), mean preference +10.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Add a module to an F# project | +0.0% | +0.0% | 0/1/0 |
| = Add a signature file to define public API | +0.0% | +0.0% | 0/1/0 |
| ▲ Fix broken file order causing FS0039 | +100.0% | +40.0% | 1/0/0 |
| = Review MSBuild files for anti-patterns and style issues | +0.0% | +0.0% | 0/1/0 |
⚠️ msbuild-modernization — details
Reason: Net win +0.0% (0W/1T/0L over 1 trial(s), sign test p=1.000), mean preference +0.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Modernize legacy project to SDK-style | +0.0% | +0.0% | 0/1/0 |
⚠️ msbuild-server — details
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Recommend MSBuild Server for slow CLI incremental builds | +100.0% | +40.0% | 1/0/0 |
⚠️ mtp-hot-reload — details
Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 6 errored, 1 unmatched — inconclusive (comparison errors)
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Enable hot reload when package already installed | +0.0% | +0.0% | 0/0/0 |
| = Negative: VSTest project cannot use MTP hot reload | +0.0% | +0.0% | 0/0/0 |
| = Run specific failing test with hot reload filter | +0.0% | +0.0% | 0/0/0 |
| = Suggest hot reload for failing test in MTP project (SDK 10) | +0.0% | +0.0% | 0/0/0 |
| = Suggest hot reload for failing test in MTP project (SDK 9) | +0.0% | +0.0% | 0/0/0 |
| = Suggest launchSettings.json configuration for hot reload | +0.0% | +0.0% | 0/0/0 |
| = Use dotnet run not dotnet test for hot reload | +0.0% | +0.0% | 0/0/0 |
⚠️ platform-detection — details
Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 5 errored — inconclusive (comparison errors)
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = MTP signal set in Directory.Build.props rather than the project file | +0.0% | +0.0% | 0/0/0 |
| = Microsoft.NET.Test.Sdk alongside an MTP runner property | +0.0% | +0.0% | 0/0/0 |
| = TUnit project is MTP-only | +0.0% | +0.0% | 0/0/0 |
| = global.json opts a plain xUnit v3 project into MTP on SDK 10 | +0.0% | +0.0% | 0/0/0 |
| = global.json runner outranks TestingPlatformDotnetTestSupport on SDK 10 | +0.0% | +0.0% | 0/0/0 |
⚠️ property-patterns — details
Reason: Net win +66.7% (2W/1T/0L over 3 trial(s), sign test p=0.250), mean preference +26.7% — underpowered (3 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Diagnose multi-level property hierarchy bugs | +100.0% | +40.0% | 1/0/0 |
| = Diagnose shared build property issues | +0.0% | +0.0% | 0/1/0 |
| ▲ Fix shared property configuration | +100.0% | +40.0% | 1/0/0 |
⚠️ resolve-project-references — details
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Explain misleading ResolveProjectReferences time | +100.0% | +40.0% | 1/0/0 |
⚠️ run-tests — details
Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 15 errored — inconclusive (comparison errors)
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Combine multiple filter criteria on VSTest MSTest | +0.0% | +0.0% | 0/0/0 |
| = Detect test platform from Directory.Build.props | +0.0% | +0.0% | 0/0/0 |
| = Filter MSTest tests by category on VSTest | +0.0% | +0.0% | 0/0/0 |
| = Filter NUnit tests by class name on VSTest | +0.0% | +0.0% | 0/0/0 |
| = Filter TUnit tests by class using treenode-filter | +0.0% | +0.0% | 0/0/0 |
| = Filter xUnit v3 tests by class on MTP | +0.0% | +0.0% | 0/0/0 |
| = Filter xUnit v3 tests by class pattern and trait using query filter language | +0.0% | +0.0% | 0/0/0 |
| = Filter xUnit v3 tests by trait on MTP | +0.0% | +0.0% | 0/0/0 |
| = MTP project on SDK 10 passes args directly | +0.0% | +0.0% | 0/0/0 |
| = MTP project on SDK 9 must use -- separator for args | +0.0% | +0.0% | 0/0/0 |
| = Negative test: do not use MTP syntax for a VSTest project | +0.0% | +0.0% | 0/0/0 |
| = Run tests in a VSTest MSTest project | +0.0% | +0.0% | 0/0/0 |
| = Run tests in a multi-TFM project targeting a specific framework | +0.0% | +0.0% | 0/0/0 |
| = Run tests with blame-hang on MTP project (SDK 10) | +0.0% | +0.0% | 0/0/0 |
| = Run tests with trx reporting on MTP project (SDK 9) | +0.0% | +0.0% | 0/0/0 |
⚠️ target-authoring — details
Reason: Net win +33.3% (1W/2T/0L over 3 trial(s), sign test p=0.500), mean preference +13.3% — underpowered (3 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Diagnose broken SDK target chain across files | +0.0% | +0.0% | 0/1/0 |
| ▲ Diagnose custom target build regression | +100.0% | +40.0% | 1/0/0 |
| = Fix custom target anti-patterns | +0.0% | +0.0% | 0/1/0 |
⚠️ test-anti-patterns — details
Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 8 errored — inconclusive (comparison errors)
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Audit a pytest suite using Python-specific anti-pattern markers | +0.0% | +0.0% | 0/0/0 |
| = Detect coverage-touching pattern across a service facade | +0.0% | +0.0% | 0/0/0 |
| = Detect duplicated tests and magic values | +0.0% | +0.0% | 0/0/0 |
| = Detect flakiness indicators and test coupling | +0.0% | +0.0% | 0/0/0 |
| = Detect mixed severity anti-patterns in repository service tests | +0.0% | +0.0% | 0/0/0 |
| = Detect self-referential assertions in round-trip and identity tests | +0.0% | +0.0% | 0/0/0 |
| = Recognize well-written tests without inventing false positives | +0.0% | +0.0% | 0/0/0 |
| = Separate false-confidence assertions from cosmetic ones | +0.0% | +0.0% | 0/0/0 |
⚠️ test-gap-analysis — details
Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 6 errored — inconclusive (comparison errors)
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Acknowledge well-tested code with few surviving mutations | +0.0% | +0.0% | 0/0/0 |
| = Analyse error propagation gaps in a Rust crate | +0.0% | +0.0% | 0/0/0 |
| = Decline request to write new tests from scratch | +0.0% | +0.0% | 0/0/0 |
| = Find boundary mutation gaps in tiered discount and shipping logic | +0.0% | +0.0% | 0/0/0 |
| = Find logic and null-check mutation gaps in access control code | +0.0% | +0.0% | 0/0/0 |
| = Skip trivial and generated code while tracing private call chains | +0.0% | +0.0% | 0/0/0 |
⚠️ test-smell-detection — details
Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 10 errored — inconclusive (comparison errors)
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Audit a JUnit suite using Java-specific smell markers | +0.0% | +0.0% | 0/0/0 |
| = Detect multiple test smells in order processing test suite | +0.0% | +0.0% | 0/0/0 |
| = Recognize integration tests and avoid false positives for external resources | +0.0% | +0.0% | 0/0/0 |
| = Recognize well-written tests with no significant smells | +0.0% | +0.0% | 0/0/0 |
| = Separate reasoned skips and self-documenting numbers from real smells | +0.0% | +0.0% | 0/0/0 |
⚠️ test-tagging — details
Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 9 errored — inconclusive (comparison errors)
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Accurately classify NUnit tests with misleading method names | +0.0% | +0.0% | 0/0/0 |
| = Audit test distribution without modifying files | +0.0% | +0.0% | 0/0/0 |
| = Classify Go tests where no trait attribute mechanism exists | +0.0% | +0.0% | 0/0/0 |
| = Decline request to write new tests | +0.0% | +0.0% | 0/0/0 |
| = Tag MSTest tests and verify the project still builds | +0.0% | +0.0% | 0/0/0 |
| = Tag a partially-tagged MSTest suite without duplicating existing traits | +0.0% | +0.0% | 0/0/0 |
| = Tag an untagged MSTest test suite | +0.0% | +0.0% | 0/0/0 |
| = Tag an untagged NUnit test suite | +0.0% | +0.0% | 0/0/0 |
| = Tag an untagged xUnit test suite | +0.0% | +0.0% | 0/0/0 |
⚠️ writing-mstest-tests — details
Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 16 errored — inconclusive (comparison errors)
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Configure conditional execution, retry, and cleanup | +0.0% | +0.0% | 0/0/0 |
| = Configure test parallelization and MSTest.Sdk project | +0.0% | +0.0% | 0/0/0 |
| = Fix swapped Assert.AreEqual arguments | +0.0% | +0.0% | 0/0/0 |
| = Modernize legacy test patterns | +0.0% | +0.0% | 0/0/0 |
| = Replace ExpectedException with Assert.Throws | +0.0% | +0.0% | 0/0/0 |
| = Replace generic IsTrue checks for null, identity, emptiness, and absence | +0.0% | +0.0% | 0/0/0 |
| = Set up test lifecycle correctly | +0.0% | +0.0% | 0/0/0 |
| = Use DynamicData with ValueTuples over object arrays | +0.0% | +0.0% | 0/0/0 |
| = Use comparison assertions for boundary testing | +0.0% | +0.0% | 0/0/0 |
| = Use proper collection assertions | +0.0% | +0.0% | 0/0/0 |
| = Use proper type assertions instead of casts | +0.0% | +0.0% | 0/0/0 |
| = Use string assertions for format validation | +0.0% | +0.0% | 0/0/0 |
| = Write async tests with cancellation | +0.0% | +0.0% | 0/0/0 |
| = Write data-driven tests for a calculator | +0.0% | +0.0% | 0/0/0 |
| = Write tests with collection, null, and reference assertions | +0.0% | +0.0% | 0/0/0 |
| = Write unit tests for a service class | +0.0% | +0.0% | 0/0/0 |
🔍 Full Results - additional metrics and failure investigation steps
▶ Sessions Visualisation -- interactive replay of all evaluation sessions
📊 Session Analytics (preview) -- aggregated metrics across evaluation sessions
|
✅ Evaluation passed for |
Bumps the github-actions-dependencies group with 2 updates in the / directory: actions/checkout and actions/setup-python.
Updates
actions/checkoutfrom 7.0.0 to 7.0.1Release notes
Sourced from actions/checkout's releases.
Changelog
Sourced from actions/checkout's changelog.
... (truncated)
Commits
Updates
actions/setup-pythonfrom 5.6.0 to 7.0.0Release notes
Sourced from actions/setup-python's releases.
... (truncated)
Commits
5fda3b9Pin SHA commits and update docs with latest versions (#1338)4ab7e95Merge pull request #1337 from actions/philip-gai/bump-actions-cache-6-2-00f3a009Remove the pip-install input (#1336)f8cf429Migrate to ESM and upgrade dependencies (#1330)54baeeaValidate and retry manifest fetch to prevent silent failures (#1332)c709277Annotation code fix (#1335)6849080remove EOL Python versions and Bumps numpy text fixture (#1333)0903b46Bump certifi from 2020.6.20 to 2024.7.4 in /tests/data (#1328)ece7cb0Fix pip cache error handling on Windows. (#1040)1d18d7aUpdate advanced-usage.md (#811)