Skip to content

Bump the github-actions-dependencies group across 1 directory with 2 updates - #963

Open
dependabot[bot] wants to merge 1 commit into
mainfrom
dependabot/github_actions/github-actions-dependencies-25c580fd63
Open

Bump the github-actions-dependencies group across 1 directory with 2 updates#963
dependabot[bot] wants to merge 1 commit into
mainfrom
dependabot/github_actions/github-actions-dependencies-25c580fd63

Conversation

@dependabot

@dependabot dependabot Bot commented on behalf of github Jul 29, 2026

Copy link
Copy Markdown
Contributor

Bumps the github-actions-dependencies group with 2 updates in the / directory: actions/checkout and actions/setup-python.

Updates actions/checkout from 7.0.0 to 7.0.1

Release notes

Sourced from actions/checkout's releases.

v7.0.1

What's Changed

Full Changelog: actions/checkout@v7...v7.0.1

Changelog

Sourced from actions/checkout's changelog.

Changelog

v7.0.1

v7.0.0

v6.0.3

v6.0.2

v6.0.1

v6.0.0

v5.0.1

v5.0.0

v4.3.1

v4.3.0

v4.2.2

v4.2.1

... (truncated)

Commits

Updates actions/setup-python from 5.6.0 to 7.0.0

Release notes

Sourced from actions/setup-python's releases.

v7.0.0

What's Changed

Enhancements

Bug Fix

Dependency Upgrade

New Contributors

Full Changelog: actions/setup-python@v6...v7.0.0

v6.3.0

What's Changed

Enhancement

Dependency update

Documentation

New Contributors

Full Changelog: actions/setup-python@v6.2.0...v6.3.0

v6.2.0

What's Changed

Dependency Upgrades

... (truncated)

Commits

@dependabot dependabot Bot added dependencies Pull requests that update a dependency file github_actions Pull requests that update GitHub Actions code labels Jul 29, 2026
@github-actions github-actions Bot added the pr-state/ready-for-eval PR is mergeable and awaiting evaluation label Jul 29, 2026
github-actions Bot added a commit that referenced this pull request Jul 29, 2026
@github-actions

Copy link
Copy Markdown
Contributor

📊 Skill Evaluation Results

10 skill(s) evaluated — 6 improved, 4 no credible improvement. A skill passes only on a credible improvement over baseline (mean preference > 0 with its 95% CI above 0); ⚠️ marks a comparison that couldn't complete (errored/unmatched trials).

Skill Result Δ Preference [95% CI] W/T/L Quality Baseline Overfit Skills Loaded
create-datadriven-aspnetcore +56.0% [+2.2%, +109.8%] 4/1/0 4.4/5 3.3/5 ✅ 0.12 ⚠️ 4/5 · 5/5 (plugin)
dotnet-maui-doctor +70.0% [+43.2%, +96.8%] 8/0/0 4.6/5 3.6/5 ✅ 0.17 8/8 · 8/8 (plugin)
maui-app-lifecycle +40.0% [+40.0%, +40.0%] 4/0/0 4.9/5 3.7/5 ✅ 0.12 4/4 · 4/4 (plugin)
maui-collectionview +30.0% [-1.8%, +61.8%] 3/1/0 5.0/5 5.0/5 ✅ 0.07 4/4 · 4/4 (plugin)
maui-data-binding +40.0% [+40.0%, +40.0%] 4/0/0 4.9/5 3.9/5 ✅ 0.14 4/4 · 4/4 (plugin)
maui-dependency-injection +20.0% [-43.6%, +83.6%] 3/0/1 5.0/5 4.6/5 ✅ 0.10 4/4 · 4/4 (plugin)
maui-safe-area +100.0% [+100.0%, +100.0%] 4/0/0 5.0/5 2.1/5 ✅ 0.08 4/4 · 4/4 (plugin)
maui-shell-navigation +55.0% [+7.3%, +102.7%] 4/0/0 5.0/5 3.6/5 ✅ 0.09 4/4 · 4/4 (plugin)
maui-theming +10.0% [-21.8%, +41.8%] 1/3/0 5.0/5 4.9/5 ✅ 0.11 4/4 · 4/4 (plugin)
optimizing-ef-core-queries -40.0% [-40.0%, -40.0%] 0/0/1 4.7/5 5.0/5 🟡 0.21 ⚠️ 1/1 · 0/1 (plugin)
ℹ️ Column legend
  • Δ Preference — mean head-to-head preference of skilled vs baseline (−100%…+100%), judged by vally compare.
  • [95% CI] — 95% confidence interval on that mean; a skill passes only when the whole interval is above 0.
  • W/T/L — wins / ties / losses across trials.
  • Quality / Baseline — mean absolute judge score 0–5 (skilled isolated vs skill-free control).
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) with its score.
  • Skills Loaded — of the scenarios that expect activation, how many actually activated / that total (plugin run shown when present); ⚠️ marks a scenario that expected activation but didn't activate.
✅ create-datadriven-aspnetcore — details

Reason: Mean preference +56.0% [95% CI 2.2%, 109.8%], win rate 80.0% (4W/1T/0L over 5 trial(s)) — credibly better

Scenario Mean preference Trials (W/T/L)
▲ Scaffold Blazor CRUD components +40.0% 1/0/0
▲ Scaffold MVC CRUD with an existing DbContext +40.0% 1/0/0
= Scaffold Minimal API child and required parent CRUD +0.0% 0/1/0
▲ Scaffold Minimal API endpoints with OpenAPI +100.0% 1/0/0
▲ Scaffold Razor Pages CRUD with EF Core and SQLite +100.0% 1/0/0
✅ dotnet-maui-doctor — details

Reason: Mean preference +70.0% [95% CI 43.2%, 96.8%], win rate 100.0% (8W/0T/0L over 8 trial(s)) — credibly better

Scenario Mean preference Trials (W/T/L)
▲ Determine required Android SDK packages for specific .NET version +40.0% 1/0/0
▲ Diagnose non-Microsoft JDK causing build failure +100.0% 1/0/0
▲ Fix stale MAUI workloads after SDK update +100.0% 1/0/0
▲ Guardrail against workload update and repair +100.0% 1/0/0
▲ Plan Linux MAUI environment for Android +40.0% 1/0/0
▲ Plan complete MAUI setup on Windows +40.0% 1/0/0
▲ Plan macOS MAUI setup with Xcode +100.0% 1/0/0
▲ Prevent incorrect JAVA_HOME configuration +40.0% 1/0/0
✅ maui-app-lifecycle — details

Reason: Mean preference +40.0% [95% CI 40.0%, 40.0%], win rate 100.0% (4W/0T/0L over 4 trial(s)) — credibly better

Scenario Mean preference Trials (W/T/L)
▲ Avoid legacy Xamarin.Forms lifecycle methods +40.0% 1/0/0
▲ Platform-specific lifecycle mapping +40.0% 1/0/0
▲ Save and restore state on background +40.0% 1/0/0
▲ Window lifecycle event subscription +40.0% 1/0/0
❌ maui-collectionview — details

Reason: Mean preference +30.0% [95% CI -1.8%, 61.8%], win rate 75.0% (3W/1T/0L over 4 trial(s)) — not credible (95% CI includes 0)

Scenario Mean preference Trials (W/T/L)
▲ Avoid ListView and ViewCell mistakes +40.0% 1/0/0
▲ Basic CollectionView with data binding and DataTemplate +40.0% 1/0/0
▲ Grid layout with CollectionView +40.0% 1/0/0
= Selection and pull-to-refresh with CollectionView +0.0% 0/1/0
✅ maui-data-binding — details

Reason: Mean preference +40.0% [95% CI 40.0%, 40.0%], win rate 100.0% (4W/0T/0L over 4 trial(s)) — credibly better

Scenario Mean preference Trials (W/T/L)
▲ Create and use an IValueConverter +40.0% 1/0/0
▲ Diagnose missing BindingContext and non-compiled bindings +40.0% 1/0/0
▲ Implement MVVM ViewModel with ObservableObject +40.0% 1/0/0
▲ Set up compiled bindings with x:DataType on a page +40.0% 1/0/0
❌ maui-dependency-injection — details

Reason: Mean preference +20.0% [95% CI -43.6%, 83.6%], win rate 75.0% (3W/0T/1L over 4 trial(s)) — not credible (95% CI includes 0)

Scenario Mean preference Trials (W/T/L)
▲ Avoid AddScoped pitfall in MAUI +40.0% 1/0/0
▲ Platform-specific service registration with fallback +40.0% 1/0/0
▲ Register services with correct lifetimes in MauiProgram.cs +40.0% 1/0/0
▼ Shell navigation auto-resolves DI-registered pages -40.0% 0/0/1
✅ maui-safe-area — details

Reason: Mean preference +100.0% [95% CI 100.0%, 100.0%], win rate 100.0% (4W/0T/0L over 4 trial(s)) — credibly better

Scenario Mean preference Trials (W/T/L)
▲ Avoid deprecated UseSafeArea in new .NET 10 project +100.0% 1/0/0
▲ Edge-to-edge layout with SafeAreaEdges +100.0% 1/0/0
▲ Handle notch and status bar safe areas on iOS +100.0% 1/0/0
▲ Keyboard avoidance with safe area for chat UI +100.0% 1/0/0
✅ maui-shell-navigation — details

Reason: Mean preference +55.0% [95% CI 7.3%, 102.7%], win rate 100.0% (4W/0T/0L over 4 trial(s)) — credibly better

Scenario Mean preference Trials (W/T/L)
▲ Diagnose common Shell navigation mistakes +40.0% 1/0/0
▲ Handle back navigation and unsaved changes guard +100.0% 1/0/0
▲ Navigate with parameters using GoToAsync +40.0% 1/0/0
▲ Set up Shell navigation with tabs and flyout +40.0% 1/0/0
❌ maui-theming — details

Reason: Mean preference +10.0% [95% CI -21.8%, 41.8%], win rate 25.0% (1W/3T/0L over 4 trial(s)) — not credible (95% CI includes 0)

Scenario Mean preference Trials (W/T/L)
= Add light/dark mode support using AppThemeBinding +0.0% 0/1/0
= Avoid common theming mistakes +0.0% 0/1/0
▲ Create custom themes with ResourceDictionary switching +40.0% 1/0/0
= Detect and respond to system theme changes +0.0% 0/1/0
❌ optimizing-ef-core-queries — details

Reason: Mean preference -40.0% [95% CI -40.0%, -40.0%], win rate 0.0% (0W/0T/1L over 1 trial(s)) — no improvement

Scenario Mean preference Trials (W/T/L)
▼ Optimize bulk operations with EF Core 7+ ExecuteUpdate and ExecuteDelete -40.0% 0/0/1

🔍 Full Results - additional metrics and failure investigation steps

To investigate failures, paste this to your AI coding agent:

For PR 963 in dotnet/skills, download eval artifacts with gh run download 30463860582 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/b6213a8ba92547bd7aee8a095cfa9e93a24297ae/eng/vally-adapter/InvestigatingResults.md and follow it to analyze the results.json files. Diagnose each failure, suggest fixes to the eval.yaml and skill content, and tell me what to fix first.

▶ Sessions Visualisation -- interactive replay of all evaluation sessions
📊 Session Analytics (preview) -- aggregated metrics across evaluation sessions

@github-actions github-actions Bot added waiting-on-review PR state label and removed pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Jul 29, 2026
@github-actions

Copy link
Copy Markdown
Contributor

✅ Evaluation passed for b6213a8. cc @AbhitejJohn @JanKrivanek — please review.

@dependabot
dependabot Bot force-pushed the dependabot/github_actions/github-actions-dependencies-25c580fd63 branch from b6213a8 to 63cfc57 Compare July 30, 2026 09:59
@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation and removed waiting-on-review PR state label labels Jul 30, 2026
@github-actions

Copy link
Copy Markdown
Contributor

❌ Evaluation ran but produced no results.

The evaluate job completed but no results.json verdicts were generated. This is usually a transient infrastructure failure — commonly an LLM-session auth error (Session was not created with authentication info or custom provider) — not a problem with your skill. Check the workflow run logs, then re-post /evaluate to try again.

@github-actions github-actions Bot added waiting-on-review PR state label and removed pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Jul 30, 2026
@github-actions

github-actions Bot commented Aug 2, 2026

Copy link
Copy Markdown
Contributor

✅ Evaluation passed for 63cfc57. cc @AbhitejJohn @JanKrivanek — please review.

@dependabot
dependabot Bot force-pushed the dependabot/github_actions/github-actions-dependencies-25c580fd63 branch from 63cfc57 to fe76869 Compare August 4, 2026 22:30
@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation and removed waiting-on-review PR state label labels Aug 4, 2026
github-actions Bot added a commit that referenced this pull request Aug 4, 2026
@github-actions

github-actions Bot commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

📊 Skill Evaluation Results

27 skill(s) evaluated — ✅ 2 improved, ❌ 0 no credible change, 🔻 0 regressed.

⚠️ 25 could not be judged: 25 underpowered — the eval has fewer trials than any result needs to reach p ≤ 0.05, so no verdict was possible. This is the eval's size, not a skill regression; fix it by adding scenarios or raising defaults.runs.

A skill passes only on a credible net win over baseline: more wins than losses, by an exact one-sided sign test at p ≤ 0.05.

Skill Result Net win p Δ Pref W/T/L Quality Baseline Overfit Skills Loaded
author-component +100.0% 0.031 +52.0% 5/0/0 3.9/5 3.3/5 ✅ 0.15 5/5 · 5/5 (plugin)
binlog-failure-analysis ⚠️ +0.0% 1.000 +0.0% 0/1/0 5.0/5 5.0/5 ✅ 0.08 1/1 · 1/1 (plugin)
binlog-generation ⚠️ +100.0% 0.125 +80.0% 3/0/0 3.3/5 1.8/5 🟡 0.37 3/3 · 3/3 (plugin)
build-parallelism ⚠️ +100.0% 0.500 +40.0% 1/0/0 4.6/5 4.6/5 🟡 0.26 1/1 · 1/1 (plugin)
build-perf-baseline ⚠️ +100.0% 0.500 +40.0% 1/0/0 2.2/5 3.1/5 🟡 0.24 1/1 · 1/1 (plugin)
build-perf-diagnostics ⚠️ +0.0% 1.000 +0.0% 0/1/0 5.0/5 5.0/5 🟡 0.28 ⚠️ 0/1 · 1/1 (plugin)
check-bin-obj-clash ⚠️ +0.0% 1.000 +0.0% 0/1/0 5.0/5 4.7/5 🟡 0.22 1/1 · 1/1 (plugin)
collect-user-input ⚠️ +50.0% 0.500 +20.0% 1/1/0 4.6/5 4.7/5 ✅ 0.07 ⚠️ 1/2 · 2/2 (plugin)
configure-auth ⚠️ +50.0% 0.500 +20.0% 1/1/0 5.0/5 2.9/5 ✅ 0.15 2/2 · 2/2 (plugin)
coordinate-components ⚠️ +50.0% 0.500 +50.0% 1/1/0 3.2/5 2.3/5 🟡 0.26 2/2 · 2/2 (plugin)
create-blazor-project ⚠️ +100.0% 0.125 +60.0% 3/0/0 3.0/5 2.6/5 ✅ 0.17 ⚠️ 3/3 · 2/3 (plugin)
directory-build-organization ⚠️ +0.0% 1.000 +0.0% 0/1/0 5.0/5 4.2/5 🟡 0.21 1/1 · 1/1 (plugin)
eval-performance ⚠️ +100.0% 0.500 +40.0% 1/0/0 4.7/5 4.4/5 ✅ 0.11 1/1 · 1/1 (plugin)
extension-points ⚠️ +100.0% 0.125 +40.0% 3/0/0 4.7/5 4.0/5 ✅ 0.07 3/3 · 3/3 (plugin)
fetch-and-send-data ⚠️ +100.0% 0.250 +40.0% 2/0/0 4.8/5 3.1/5 ✅ 0.16 ⚠️ 2/2 · 1/2 (plugin)
including-generated-files ⚠️ +100.0% 0.500 +40.0% 1/0/0 5.0/5 2.5/5 🟡 0.23 1/1 · 1/1 (plugin)
incremental-build ⚠️ +100.0% 0.500 +40.0% 1/0/0 4.6/5 4.2/5 ✅ 0.16 1/1 · 1/1 (plugin)
item-management ⚠️ +0.0% 0.750 +0.0% 1/1/1 4.9/5 4.6/5 ✅ 0.15 ⚠️ 3/3 · 2/3 (plugin)
msbuild-antipatterns ⚠️ +0.0% 1.000 +0.0% 0/4/0 4.9/5 4.9/5 ✅ 0.07 ⚠️ 1/4 · 1/4 (plugin)
msbuild-modernization ⚠️ +100.0% 0.500 +40.0% 1/0/0 5.0/5 5.0/5 ✅ 0.06 1/1 · 1/1 (plugin)
msbuild-server ⚠️ +100.0% 0.500 +40.0% 1/0/0 5.0/5 3.1/5 🟡 0.32 1/1 · 1/1 (plugin)
plan-ui-change +100.0% 0.031 +64.0% 5/0/0 2.9/5 1.8/5 🟡 0.24 5/5 · 5/5 (plugin)
property-patterns ⚠️ -33.3% 0.500 -13.3% 0/2/1 4.7/5 4.9/5 ✅ 0.08 ⚠️ 2/3 · 2/3 (plugin)
resolve-project-references ⚠️ +0.0% 1.000 +0.0% 0/1/0 5.0/5 3.8/5 ✅ 0.16 1/1 · 1/1 (plugin)
support-prerendering ⚠️ +100.0% 0.250 +70.0% 2/0/0 4.0/5 1.8/5 🟡 0.24 2/2 · 2/2 (plugin)
target-authoring ⚠️ +100.0% 0.125 +40.0% 3/0/0 4.7/5 4.4/5 🟡 0.20 ⚠️ 3/3 · 2/3 (plugin)
use-js-interop ⚠️ +100.0% 0.063 +40.0% 4/0/0 1.3/5 1.4/5 ✅ 0.18 4/4 · 4/4 (plugin)
ℹ️ Column legend
  • Net win(wins − losses) / trials for skilled vs baseline, judged head-to-head by vally compare. This is the effect the gate decides on.
  • p — one-sided exact sign test over the discordant (non-tie) trials. A skill passes only at p ≤ 0.05, which needs at least 5 winning trials.
  • Δ Pref — the same comparison weighted by how decisive each win was (much-better ±100%, slightly-better ±40%). Reported for triage only: weighting the statistic by magnitude made a skill fail for winning harder, which is why the gate deliberately ignores this column.
  • W/T/L — wins / ties / losses across trials.
  • ⚠️ — the gate withheld a verdict. Either the eval has fewer trials than any result needs to reach p ≤ 0.05 (underpowered — the skill was never actually measured, so this is not a regression; add scenarios or raise defaults.runs), or the comparison didn't complete.
  • 🔻 — a credible regression: the losses themselves clear the same bar the gate uses for wins.
  • Quality / Baseline — mean absolute judge score 0–5 (skilled isolated vs skill-free control).
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) with its score.
  • Skills Loaded — of the scenarios that expect activation, how many actually activated / that total (plugin run shown when present); ⚠️ marks a scenario that expected activation but didn't activate.
⚠️ binlog-failure-analysis — details

Reason: Net win +0.0% (0W/1T/0L over 1 trial(s), sign test p=1.000), mean preference +0.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
= Diagnose build failures from binlog only (no source files) +0.0% +0.0% 0/1/0
⚠️ binlog-generation — details

Reason: Net win +100.0% (3W/0T/0L over 3 trial(s), sign test p=0.125), mean preference +80.0% — underpowered (3 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Build multiple configurations with unique binlogs +100.0% +100.0% 1/0/0
▲ Build project with /bl flag +100.0% +100.0% 1/0/0
▲ Build with /bl in PowerShell +100.0% +40.0% 1/0/0
⚠️ build-parallelism — details

Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Analyze build parallelism bottlenecks +100.0% +40.0% 1/0/0
⚠️ build-perf-baseline — details

Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Establish build performance baseline and recommend optimizations +100.0% +40.0% 1/0/0
⚠️ build-perf-diagnostics — details

Reason: Net win +0.0% (0W/1T/0L over 1 trial(s), sign test p=1.000), mean preference +0.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
= Diagnose slow build for a small project +0.0% +0.0% 0/1/0
⚠️ check-bin-obj-clash — details

Reason: Net win +0.0% (0W/1T/0L over 1 trial(s), sign test p=1.000), mean preference +0.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
= Diagnose bin/obj output path clashes +0.0% +0.0% 0/1/0
⚠️ collect-user-input — details

Reason: Net win +50.0% (1W/1T/0L over 2 trial(s), sign test p=0.500), mean preference +20.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Event registration with custom validation +100.0% +40.0% 1/0/0
= Multi-step booking form with cross-field validation +0.0% +0.0% 0/1/0
⚠️ configure-auth — details

Reason: Net win +50.0% (1W/1T/0L over 2 trial(s), sign test p=0.500), mean preference +20.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
= Login and account management in a globally interactive app +0.0% +0.0% 0/1/0
▲ Multi-tier app with WebAssembly auth +100.0% +40.0% 1/0/0
⚠️ coordinate-components — details

Reason: Net win +50.0% (1W/1T/0L over 2 trial(s), sign test p=0.500), mean preference +50.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Multi-tenant notification hub with cross-component fan-out +100.0% +100.0% 1/0/0
= Warehouse dashboard with site selector and live stock alerts +0.0% +0.0% 0/1/0
⚠️ create-blazor-project — details

Reason: Net win +100.0% (3W/0T/0L over 3 trial(s), sign test p=0.125), mean preference +60.0% — underpowered (3 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Global logistics tracking for worldwide users +100.0% +100.0% 1/0/0
▲ Recipe community with interactive ratings on static pages +100.0% +40.0% 1/0/0
▲ University course catalog with enrollment form +100.0% +40.0% 1/0/0
⚠️ directory-build-organization — details

Reason: Net win +0.0% (0W/1T/0L over 1 trial(s), sign test p=1.000), mean preference +0.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
= Organize build infrastructure for a multi-project repo +0.0% +0.0% 0/1/0
⚠️ eval-performance — details

Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Analyze MSBuild evaluation performance issues +100.0% +40.0% 1/0/0
⚠️ extension-points — details

Reason: Net win +100.0% (3W/0T/0L over 3 trial(s), sign test p=0.125), mean preference +40.0% — underpowered (3 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Diagnose NuGet package and repo extension conflicts +100.0% +40.0% 1/0/0
▲ Diagnose build extension point failures +100.0% +40.0% 1/0/0
▲ Fix extension point anti-patterns +100.0% +40.0% 1/0/0
⚠️ fetch-and-send-data — details

Reason: Net win +100.0% (2W/0T/0L over 2 trial(s), sign test p=0.250), mean preference +40.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Real-time shipment tracker with Auto interactivity +100.0% +40.0% 1/0/0
▲ Recipe browser with resilient data loading +100.0% +40.0% 1/0/0
⚠️ including-generated-files — details

Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Diagnose generated file inclusion failure +100.0% +40.0% 1/0/0
⚠️ incremental-build — details

Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Analyze incremental build issues +100.0% +40.0% 1/0/0
⚠️ item-management — details

Reason: Net win +0.0% (1W/1T/1L over 3 trial(s), sign test p=0.750), mean preference +0.0% — underpowered (3 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▼ Diagnose cascading item and batching bugs in code generation pipeline -100.0% -40.0% 0/0/1
= Diagnose item group and batching issues +0.0% +0.0% 0/1/0
▲ Fix item management anti-patterns +100.0% +40.0% 1/0/0
⚠️ msbuild-antipatterns — details

Reason: Net win +0.0% (0W/4T/0L over 4 trial(s), sign test p=1.000), mean preference +0.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
= Add a module to an F# project +0.0% +0.0% 0/1/0
= Add a signature file to define public API +0.0% +0.0% 0/1/0
= Fix broken file order causing FS0039 +0.0% +0.0% 0/1/0
= Review MSBuild files for anti-patterns and style issues +0.0% +0.0% 0/1/0
⚠️ msbuild-modernization — details

Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Modernize legacy project to SDK-style +100.0% +40.0% 1/0/0
⚠️ msbuild-server — details

Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Recommend MSBuild Server for slow CLI incremental builds +100.0% +40.0% 1/0/0
⚠️ property-patterns — details

Reason: Net win -33.3% (0W/2T/1L over 3 trial(s), sign test p=0.500), mean preference -13.3% — underpowered (3 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
= Diagnose multi-level property hierarchy bugs +0.0% +0.0% 0/1/0
= Diagnose shared build property issues +0.0% +0.0% 0/1/0
▼ Fix shared property configuration -100.0% -40.0% 0/0/1
⚠️ resolve-project-references — details

Reason: Net win +0.0% (0W/1T/0L over 1 trial(s), sign test p=1.000), mean preference +0.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
= Explain misleading ResolveProjectReferences time +0.0% +0.0% 0/1/0
⚠️ support-prerendering — details

Reason: Net win +100.0% (2W/0T/0L over 2 trial(s), sign test p=0.250), mean preference +70.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Equipment inventory loaded once +100.0% +100.0% 1/0/0
▲ Notifications page with live polling +100.0% +40.0% 1/0/0
⚠️ target-authoring — details

Reason: Net win +100.0% (3W/0T/0L over 3 trial(s), sign test p=0.125), mean preference +40.0% — underpowered (3 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Diagnose broken SDK target chain across files +100.0% +40.0% 1/0/0
▲ Diagnose custom target build regression +100.0% +40.0% 1/0/0
▲ Fix custom target anti-patterns +100.0% +40.0% 1/0/0
⚠️ use-js-interop — details

Reason: Net win +100.0% (4W/0T/0L over 4 trial(s), sign test p=0.063), mean preference +40.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Auto-saving notepad that survives page reloads +100.0% +40.0% 1/0/0
▲ Infinite scroll list using IntersectionObserver +100.0% +40.0% 1/0/0
▲ Responsive layout that adapts to screen size +100.0% +40.0% 1/0/0
▲ User activity tracker that detects idle timeout +100.0% +40.0% 1/0/0

Per-scenario details for 2 skill(s) were omitted to keep this comment under GitHub's 65,536-character limit — open the job's step summary or Full Results for the complete breakdown.

🔍 Full Results - additional metrics and failure investigation steps

▶ Sessions Visualisation -- interactive replay of all evaluation sessions
📊 Session Analytics (preview) -- aggregated metrics across evaluation sessions

@github-actions github-actions Bot added waiting-on-review PR state label and removed pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Aug 4, 2026
…updates

Bumps the github-actions-dependencies group with 2 updates in the / directory: [actions/checkout](https://github.com/actions/checkout) and [actions/setup-python](https://github.com/actions/setup-python).


Updates `actions/checkout` from 7.0.0 to 7.0.1
- [Release notes](https://github.com/actions/checkout/releases)
- [Changelog](https://github.com/actions/checkout/blob/main/CHANGELOG.md)
- [Commits](actions/checkout@v7...3d3c42e)

Updates `actions/setup-python` from 5.6.0 to 7.0.0
- [Release notes](https://github.com/actions/setup-python/releases)
- [Commits](actions/setup-python@a26af69...5fda3b9)

---
updated-dependencies:
- dependency-name: actions/checkout
  dependency-version: 7.0.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: github-actions-dependencies
- dependency-name: actions/setup-python
  dependency-version: 7.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
  dependency-group: github-actions-dependencies
...

Signed-off-by: dependabot[bot] <support@github.com>
@dependabot
dependabot Bot force-pushed the dependabot/github_actions/github-actions-dependencies-25c580fd63 branch from fe76869 to adbbc11 Compare August 5, 2026 19:02
@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation and removed waiting-on-review PR state label labels Aug 5, 2026
github-actions Bot added a commit that referenced this pull request Aug 5, 2026
@github-actions

github-actions Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

📊 Skill Evaluation Results

36 skill(s) evaluated — ✅ 0 improved, ❌ 0 no credible change, 🔻 0 regressed.

⚠️ 36 could not be judged: 18 underpowered — the eval has fewer trials than any result needs to reach p ≤ 0.05, so no verdict was possible. This is the eval's size, not a skill regression; fix it by adding scenarios or raising defaults.runs; 18 inconclusive — the comparison didn't complete (errored, unmatched, or self-contradictory trials).

A skill passes only on a credible net win over baseline: more wins than losses, by an exact one-sided sign test at p ≤ 0.05.

Skill Result Net win p Δ Pref W/T/L Quality Baseline Overfit Skills Loaded
assertion-quality ⚠️ +0.0% 1.000 +0.0% 0/0/0 4.9/5 4.6/5 ✅ 0.07 6/6 · 6/6 (plugin)
binlog-failure-analysis ⚠️ +0.0% 1.000 +0.0% 0/1/0 5.0/5 5.0/5 ✅ 0.08 1/1 · 1/1 (plugin)
binlog-generation ⚠️ +100.0% 0.125 +100.0% 3/0/0 3.3/5 2.9/5 🟡 0.41 3/3 · 3/3 (plugin)
build-parallelism ⚠️ +100.0% 0.500 +40.0% 1/0/0 4.6/5 4.6/5 🟡 0.26 1/1 · 1/1 (plugin)
build-perf-baseline ⚠️ +100.0% 0.500 +40.0% 1/0/0 5.0/5 1.9/5 🟡 0.36 1/1 · 1/1 (plugin)
build-perf-diagnostics ⚠️ +100.0% 0.500 +40.0% 1/0/0 5.0/5 4.4/5 🟡 0.30 1/1 · 1/1 (plugin)
check-bin-obj-clash ⚠️ -100.0% 0.500 -40.0% 0/0/1 0.6/5 5.0/5 🟡 0.31 1/1 · 1/1 (plugin)
code-testing-agent ⚠️ +0.0% 1.000 +0.0% 0/0/0 4.0/5 3.9/5 ✅ 0.19 ⚠️ 1/4 · 1/4 (plugin)
coverage-analysis ⚠️ +0.0% 1.000 +0.0% 0/0/0 4.7/5 3.3/5 ✅ 0.10 ⚠️ 6/7 · 6/7 (plugin)
crap-score ⚠️ +0.0% 1.000 +0.0% 0/0/0 3.8/5 3.8/5 ✅ 0.10 6/6 · 6/6 (plugin)
detect-static-dependencies ⚠️ +0.0% 1.000 +0.0% 0/0/0 4.4/5 4.4/5 ✅ 0.13 ⚠️ 6/7 · 6/7 (plugin)
directory-build-organization ⚠️ +100.0% 0.500 +40.0% 1/0/0 5.0/5 4.2/5 🟡 0.20 1/1 · 1/1 (plugin)
eval-performance ⚠️ +100.0% 0.500 +40.0% 1/0/0 4.7/5 4.4/5 ✅ 0.20 1/1 · 1/1 (plugin)
extension-points ⚠️ +33.3% 0.500 +13.3% 2/0/1 4.7/5 3.8/5 ✅ 0.07 3/3 · 3/3 (plugin)
filter-syntax ⚠️ +0.0% 1.000 +0.0% 0/0/0 3.3/5 2.8/5 🟡 0.22 ⚠️ 0/5 · 4/5 (plugin)
find-untested-sources ⚠️ +0.0% 1.000 +0.0% 0/0/0 4.6/5 3.8/5 🟡 0.37 ⚠️ 6/6 · 4/6 (plugin)
generate-testability-wrappers ⚠️ +0.0% 1.000 +0.0% 0/0/0 4.2/5 3.2/5 ✅ 0.16 ⚠️ 4/4 · 3/4 (plugin)
grade-tests ⚠️ +0.0% 1.000 +0.0% 0/0/0 4.9/5 2.4/5 🟡 0.38 6/6 · 6/6 (plugin)
including-generated-files ⚠️ +100.0% 0.500 +40.0% 1/0/0 5.0/5 2.5/5 🟡 0.25 1/1 · 1/1 (plugin)
incremental-build ⚠️ +100.0% 0.500 +40.0% 1/0/0 4.6/5 4.2/5 ✅ 0.14 1/1 · 1/1 (plugin)
item-management ⚠️ +100.0% 0.125 +40.0% 3/0/0 5.0/5 4.6/5 ✅ 0.16 3/3 · 3/3 (plugin)
migrate-static-to-wrapper ⚠️ +0.0% 1.000 +0.0% 0/0/0 4.8/5 4.1/5 ✅ 0.09 ⚠️ 5/6 · 4/6 (plugin)
msbuild-antipatterns ⚠️ +25.0% 0.500 +10.0% 1/3/0 4.9/5 4.9/5 ✅ 0.07 ⚠️ 1/4 · 1/4 (plugin)
msbuild-modernization ⚠️ +0.0% 1.000 +0.0% 0/1/0 5.0/5 5.0/5 ✅ 0.06 1/1 · 1/1 (plugin)
msbuild-server ⚠️ +100.0% 0.500 +40.0% 1/0/0 5.0/5 3.8/5 🟡 0.23 1/1 · 1/1 (plugin)
mtp-hot-reload ⚠️ +0.0% 1.000 +0.0% 0/0/0 4.2/5 1.0/5 ✅ 0.11 7/7 · 7/7 (plugin)
platform-detection ⚠️ +0.0% 1.000 +0.0% 0/0/0 3.6/5 4.2/5 🟡 0.23 ⚠️ 0/5 · 0/5 (plugin)
property-patterns ⚠️ +66.7% 0.250 +26.7% 2/1/0 4.9/5 4.7/5 ✅ 0.08 3/3 · 3/3 (plugin)
resolve-project-references ⚠️ +100.0% 0.500 +40.0% 1/0/0 5.0/5 3.8/5 ✅ 0.14 1/1 · 1/1 (plugin)
run-tests ⚠️ +0.0% 1.000 +0.0% 0/0/0 5.0/5 3.3/5 ✅ 0.08 ⚠️ 15/15 · 12/15 (plugin)
target-authoring ⚠️ +33.3% 0.500 +13.3% 1/2/0 4.7/5 4.4/5 ✅ 0.07 3/3 · 3/3 (plugin)
test-anti-patterns ⚠️ +0.0% 1.000 +0.0% 0/0/0 4.9/5 4.4/5 ✅ 0.19 ⚠️ 7/8 · 7/8 (plugin)
test-gap-analysis ⚠️ +0.0% 1.000 +0.0% 0/0/0 4.7/5 4.7/5 ✅ 0.10 5/5 · 5/5 (plugin)
test-smell-detection ⚠️ +0.0% 1.000 +0.0% 0/0/0 4.8/5 4.7/5 🟡 0.35 5/5 · 5/5 (plugin)
test-tagging ⚠️ +0.0% 1.000 +0.0% 0/0/0 4.6/5 4.0/5 🟡 0.26 ⚠️ 6/8 · 7/8 (plugin)
writing-mstest-tests ⚠️ +0.0% 1.000 +0.0% 0/0/0 4.6/5 3.5/5 🟡 0.28 ⚠️ 15/16 · 12/16 (plugin)
ℹ️ Column legend
  • Net win(wins − losses) / trials for skilled vs baseline, judged head-to-head by vally compare. This is the effect the gate decides on.
  • p — one-sided exact sign test over the discordant (non-tie) trials. A skill passes only at p ≤ 0.05, which needs at least 5 winning trials.
  • Δ Pref — the same comparison weighted by how decisive each win was (much-better ±100%, slightly-better ±40%). Reported for triage only: weighting the statistic by magnitude made a skill fail for winning harder, which is why the gate deliberately ignores this column.
  • W/T/L — wins / ties / losses across trials.
  • ⚠️ — the gate withheld a verdict. Either the eval has fewer trials than any result needs to reach p ≤ 0.05 (underpowered — the skill was never actually measured, so this is not a regression; add scenarios or raise defaults.runs), or the comparison didn't complete.
  • 🔻 — a credible regression: the losses themselves clear the same bar the gate uses for wins.
  • Quality / Baseline — mean absolute judge score 0–5 (skilled isolated vs skill-free control).
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) with its score.
  • Skills Loaded — of the scenarios that expect activation, how many actually activated / that total (plugin run shown when present); ⚠️ marks a scenario that expected activation but didn't activate.
⚠️ assertion-quality — details

Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 12 errored — inconclusive (comparison errors)

Scenario Net win Δ Pref Trials (W/T/L)
= Distinguish weak and meaningful assertions in a pytest suite +0.0% +0.0% 0/0/0
= Flag assertion-free tests and trivial-only assertions +0.0% +0.0% 0/0/0
= Identify low assertion diversity in equality-dominated test suite +0.0% +0.0% 0/0/0
= Identify self-referential assertions in identity and round-trip tests +0.0% +0.0% 0/0/0
= Judge assertion strength in a shallow Jest suite +0.0% +0.0% 0/0/0
= Recognize well-diversified assertion usage +0.0% +0.0% 0/0/0
⚠️ binlog-failure-analysis — details

Reason: Net win +0.0% (0W/1T/0L over 1 trial(s), sign test p=1.000), mean preference +0.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
= Diagnose build failures from binlog only (no source files) +0.0% +0.0% 0/1/0
⚠️ binlog-generation — details

Reason: Net win +100.0% (3W/0T/0L over 3 trial(s), sign test p=0.125), mean preference +100.0% — underpowered (3 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Build multiple configurations with unique binlogs +100.0% +100.0% 1/0/0
▲ Build project with /bl flag +100.0% +100.0% 1/0/0
▲ Build with /bl in PowerShell +100.0% +100.0% 1/0/0
⚠️ build-parallelism — details

Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Analyze build parallelism bottlenecks +100.0% +40.0% 1/0/0
⚠️ build-perf-baseline — details

Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Establish build performance baseline and recommend optimizations +100.0% +40.0% 1/0/0
⚠️ build-perf-diagnostics — details

Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Diagnose slow build for a small project +100.0% +40.0% 1/0/0
⚠️ check-bin-obj-clash — details

Reason: Net win -100.0% (0W/0T/1L over 1 trial(s), sign test p=0.500), mean preference -40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▼ Diagnose bin/obj output path clashes -100.0% -40.0% 0/0/1
⚠️ code-testing-agent — details

Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 10 errored — inconclusive (comparison errors)

Scenario Net win Δ Pref Trials (W/T/L)
= Diagnose failing tests without generating a new suite +0.0% +0.0% 0/0/0
= Does not revert a gutted-looking workspace (workspace integrity) +0.0% +0.0% 0/0/0
= Extend an existing suite to the untested method only +0.0% +0.0% 0/0/0
= Generate Vitest tests for the shopping-cart library (TypeScript polyglot) +0.0% +0.0% 0/0/0
= Keep a single-function request proportional +0.0% +0.0% 0/0/0
⚠️ coverage-analysis — details

Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 21 errored — inconclusive (comparison errors)

Scenario Net win Δ Pref Trials (W/T/L)
= Account for a gap spread across several members +0.0% +0.0% 0/0/0
= Analyse a CI Cobertura report without re-running tests or installing tools +0.0% +0.0% 0/0/0
= Coverage plateau diagnosis +0.0% +0.0% 0/0/0
= Distinguish branch coverage from line coverage +0.0% +0.0% 0/0/0
= Project-wide coverage analysis with existing Cobertura data +0.0% +0.0% 0/0/0
= Refactoring safety assessment from coverage data +0.0% +0.0% 0/0/0
= Run coverage from scratch without existing data +0.0% +0.0% 0/0/0
⚠️ crap-score — details

Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 5 errored, 1 unmatched — inconclusive (comparison errors)

Scenario Net win Δ Pref Trials (W/T/L)
= Calculate CRAP score for a single method with partial coverage +0.0% +0.0% 0/0/0
= Generate coverage then compute CRAP score +0.0% +0.0% 0/0/0
= Identify riskiest methods across a file +0.0% +0.0% 0/0/0
= Recognize when complexity alone blocks the CRAP threshold +0.0% +0.0% 0/0/0
= Recompute complexity instead of trusting a stale source comment +0.0% +0.0% 0/0/0
= Report a fully covered method at its complexity floor +0.0% +0.0% 0/0/0
⚠️ detect-static-dependencies — details

Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 7 errored — inconclusive (comparison errors)

Scenario Net win Δ Pref Trials (W/T/L)
= Decline scan for non-C# project +0.0% +0.0% 0/0/0
= Detect statics inside lambda expressions and LINQ queries +0.0% +0.0% 0/0/0
= Detect time-related statics and recommend TimeProvider +0.0% +0.0% 0/0/0
= Exclude obj and bin directories from the scan +0.0% +0.0% 0/0/0
= Identify static dependencies in a multi-class project +0.0% +0.0% 0/0/0
= Keep one authoritative total with file line locations and no seam for pure helpers +0.0% +0.0% 0/0/0
= Verify structured report includes file count, categories, and top patterns +0.0% +0.0% 0/0/0
⚠️ directory-build-organization — details

Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Organize build infrastructure for a multi-project repo +100.0% +40.0% 1/0/0
⚠️ eval-performance — details

Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Analyze MSBuild evaluation performance issues +100.0% +40.0% 1/0/0
⚠️ extension-points — details

Reason: Net win +33.3% (2W/0T/1L over 3 trial(s), sign test p=0.500), mean preference +13.3% — underpowered (3 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Diagnose NuGet package and repo extension conflicts +100.0% +40.0% 1/0/0
▲ Diagnose build extension point failures +100.0% +40.0% 1/0/0
▼ Fix extension point anti-patterns -100.0% -40.0% 0/0/1
⚠️ filter-syntax — details

Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 5 errored — inconclusive (comparison errors)

Scenario Net win Δ Pref Trials (W/T/L)
= Filter a TUnit suite down to one class and one property value +0.0% +0.0% 0/0/0
= Filter xUnit v3 tests that do not accept the generic filter expression +0.0% +0.0% 0/0/0
= Pass a filter to a Microsoft.Testing.Platform project on the .NET 9 SDK +0.0% +0.0% 0/0/0
= Select one category and exclude another on a VSTest project +0.0% +0.0% 0/0/0
= Translate CI filter expressions after moving to xUnit v3 +0.0% +0.0% 0/0/0
⚠️ find-untested-sources — details

Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 18 errored — inconclusive (comparison errors)

Scenario Net win Δ Pref Trials (W/T/L)
= Choose the polyglot engine for a mixed C# and TypeScript repository +0.0% +0.0% 0/0/0
= Disambiguate duplicate C# type names by namespace +0.0% +0.0% 0/0/0
= Disambiguate duplicate C# types by nested test namespace +0.0% +0.0% 0/0/0
= Exclude generated sources and surface an orphan test +0.0% +0.0% 0/0/0
= Identify an unpaired TypeScript module +0.0% +0.0% 0/0/0
= Pair sources to tests across a src/tests directory split +0.0% +0.0% 0/0/0
⚠️ generate-testability-wrappers — details

Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 15 errored — inconclusive (comparison errors)

Scenario Net win Δ Pref Trials (W/T/L)
= Decline wrapper generation for already-abstracted code +0.0% +0.0% 0/0/0
= Generate TimeProvider adoption for DateTime.UtcNow +0.0% +0.0% 0/0/0
= Generate custom Environment wrapper +0.0% +0.0% 0/0/0
= Make time controllable in a library that has no DI container +0.0% +0.0% 0/0/0
= Recommend System.IO.Abstractions for file system calls +0.0% +0.0% 0/0/0
⚠️ grade-tests — details

Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 18 errored — inconclusive (comparison errors)

Scenario Net win Δ Pref Trials (W/T/L)
= Ask for a bounded list instead of grading the workspace +0.0% +0.0% 0/0/0
= Grade C# tests against available production code +0.0% +0.0% 0/0/0
= Grade Go table-driven tests without misreading the loop as branching +0.0% +0.0% 0/0/0
= Grade pytest test methods using the same rubric +0.0% +0.0% 0/0/0
= Grade tests when the production code under test is unavailable +0.0% +0.0% 0/0/0
= Keep a 62-test grading report readable as a PR comment +0.0% +0.0% 0/0/0
⚠️ including-generated-files — details

Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Diagnose generated file inclusion failure +100.0% +40.0% 1/0/0
⚠️ incremental-build — details

Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Analyze incremental build issues +100.0% +40.0% 1/0/0
⚠️ item-management — details

Reason: Net win +100.0% (3W/0T/0L over 3 trial(s), sign test p=0.125), mean preference +40.0% — underpowered (3 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Diagnose cascading item and batching bugs in code generation pipeline +100.0% +40.0% 1/0/0
▲ Diagnose item group and batching issues +100.0% +40.0% 1/0/0
▲ Fix item management anti-patterns +100.0% +40.0% 1/0/0
⚠️ migrate-static-to-wrapper — details

Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 6 errored — inconclusive (comparison errors)

Scenario Net win Δ Pref Trials (W/T/L)
= Add the required using directive and update tests with a test double +0.0% +0.0% 0/0/0
= Decline migration when wrapper does not exist yet +0.0% +0.0% 0/0/0
= Migrate DateTime.UtcNow to TimeProvider in a service class +0.0% +0.0% 0/0/0
= Migrate a static helper class without breaking its callers +0.0% +0.0% 0/0/0
= Migrate only in scoped files, leaving others untouched +0.0% +0.0% 0/0/0
= Preserve DateTimeKind when migrating to TimeProvider +0.0% +0.0% 0/0/0
⚠️ msbuild-antipatterns — details

Reason: Net win +25.0% (1W/3T/0L over 4 trial(s), sign test p=0.500), mean preference +10.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
= Add a module to an F# project +0.0% +0.0% 0/1/0
= Add a signature file to define public API +0.0% +0.0% 0/1/0
▲ Fix broken file order causing FS0039 +100.0% +40.0% 1/0/0
= Review MSBuild files for anti-patterns and style issues +0.0% +0.0% 0/1/0
⚠️ msbuild-modernization — details

Reason: Net win +0.0% (0W/1T/0L over 1 trial(s), sign test p=1.000), mean preference +0.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
= Modernize legacy project to SDK-style +0.0% +0.0% 0/1/0
⚠️ msbuild-server — details

Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Recommend MSBuild Server for slow CLI incremental builds +100.0% +40.0% 1/0/0
⚠️ mtp-hot-reload — details

Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 6 errored, 1 unmatched — inconclusive (comparison errors)

Scenario Net win Δ Pref Trials (W/T/L)
= Enable hot reload when package already installed +0.0% +0.0% 0/0/0
= Negative: VSTest project cannot use MTP hot reload +0.0% +0.0% 0/0/0
= Run specific failing test with hot reload filter +0.0% +0.0% 0/0/0
= Suggest hot reload for failing test in MTP project (SDK 10) +0.0% +0.0% 0/0/0
= Suggest hot reload for failing test in MTP project (SDK 9) +0.0% +0.0% 0/0/0
= Suggest launchSettings.json configuration for hot reload +0.0% +0.0% 0/0/0
= Use dotnet run not dotnet test for hot reload +0.0% +0.0% 0/0/0
⚠️ platform-detection — details

Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 5 errored — inconclusive (comparison errors)

Scenario Net win Δ Pref Trials (W/T/L)
= MTP signal set in Directory.Build.props rather than the project file +0.0% +0.0% 0/0/0
= Microsoft.NET.Test.Sdk alongside an MTP runner property +0.0% +0.0% 0/0/0
= TUnit project is MTP-only +0.0% +0.0% 0/0/0
= global.json opts a plain xUnit v3 project into MTP on SDK 10 +0.0% +0.0% 0/0/0
= global.json runner outranks TestingPlatformDotnetTestSupport on SDK 10 +0.0% +0.0% 0/0/0
⚠️ property-patterns — details

Reason: Net win +66.7% (2W/1T/0L over 3 trial(s), sign test p=0.250), mean preference +26.7% — underpowered (3 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Diagnose multi-level property hierarchy bugs +100.0% +40.0% 1/0/0
= Diagnose shared build property issues +0.0% +0.0% 0/1/0
▲ Fix shared property configuration +100.0% +40.0% 1/0/0
⚠️ resolve-project-references — details

Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
▲ Explain misleading ResolveProjectReferences time +100.0% +40.0% 1/0/0
⚠️ run-tests — details

Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 15 errored — inconclusive (comparison errors)

Scenario Net win Δ Pref Trials (W/T/L)
= Combine multiple filter criteria on VSTest MSTest +0.0% +0.0% 0/0/0
= Detect test platform from Directory.Build.props +0.0% +0.0% 0/0/0
= Filter MSTest tests by category on VSTest +0.0% +0.0% 0/0/0
= Filter NUnit tests by class name on VSTest +0.0% +0.0% 0/0/0
= Filter TUnit tests by class using treenode-filter +0.0% +0.0% 0/0/0
= Filter xUnit v3 tests by class on MTP +0.0% +0.0% 0/0/0
= Filter xUnit v3 tests by class pattern and trait using query filter language +0.0% +0.0% 0/0/0
= Filter xUnit v3 tests by trait on MTP +0.0% +0.0% 0/0/0
= MTP project on SDK 10 passes args directly +0.0% +0.0% 0/0/0
= MTP project on SDK 9 must use -- separator for args +0.0% +0.0% 0/0/0
= Negative test: do not use MTP syntax for a VSTest project +0.0% +0.0% 0/0/0
= Run tests in a VSTest MSTest project +0.0% +0.0% 0/0/0
= Run tests in a multi-TFM project targeting a specific framework +0.0% +0.0% 0/0/0
= Run tests with blame-hang on MTP project (SDK 10) +0.0% +0.0% 0/0/0
= Run tests with trx reporting on MTP project (SDK 9) +0.0% +0.0% 0/0/0
⚠️ target-authoring — details

Reason: Net win +33.3% (1W/2T/0L over 3 trial(s), sign test p=0.500), mean preference +13.3% — underpowered (3 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs

Scenario Net win Δ Pref Trials (W/T/L)
= Diagnose broken SDK target chain across files +0.0% +0.0% 0/1/0
▲ Diagnose custom target build regression +100.0% +40.0% 1/0/0
= Fix custom target anti-patterns +0.0% +0.0% 0/1/0
⚠️ test-anti-patterns — details

Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 8 errored — inconclusive (comparison errors)

Scenario Net win Δ Pref Trials (W/T/L)
= Audit a pytest suite using Python-specific anti-pattern markers +0.0% +0.0% 0/0/0
= Detect coverage-touching pattern across a service facade +0.0% +0.0% 0/0/0
= Detect duplicated tests and magic values +0.0% +0.0% 0/0/0
= Detect flakiness indicators and test coupling +0.0% +0.0% 0/0/0
= Detect mixed severity anti-patterns in repository service tests +0.0% +0.0% 0/0/0
= Detect self-referential assertions in round-trip and identity tests +0.0% +0.0% 0/0/0
= Recognize well-written tests without inventing false positives +0.0% +0.0% 0/0/0
= Separate false-confidence assertions from cosmetic ones +0.0% +0.0% 0/0/0
⚠️ test-gap-analysis — details

Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 6 errored — inconclusive (comparison errors)

Scenario Net win Δ Pref Trials (W/T/L)
= Acknowledge well-tested code with few surviving mutations +0.0% +0.0% 0/0/0
= Analyse error propagation gaps in a Rust crate +0.0% +0.0% 0/0/0
= Decline request to write new tests from scratch +0.0% +0.0% 0/0/0
= Find boundary mutation gaps in tiered discount and shipping logic +0.0% +0.0% 0/0/0
= Find logic and null-check mutation gaps in access control code +0.0% +0.0% 0/0/0
= Skip trivial and generated code while tracing private call chains +0.0% +0.0% 0/0/0
⚠️ test-smell-detection — details

Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 10 errored — inconclusive (comparison errors)

Scenario Net win Δ Pref Trials (W/T/L)
= Audit a JUnit suite using Java-specific smell markers +0.0% +0.0% 0/0/0
= Detect multiple test smells in order processing test suite +0.0% +0.0% 0/0/0
= Recognize integration tests and avoid false positives for external resources +0.0% +0.0% 0/0/0
= Recognize well-written tests with no significant smells +0.0% +0.0% 0/0/0
= Separate reasoned skips and self-documenting numbers from real smells +0.0% +0.0% 0/0/0
⚠️ test-tagging — details

Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 9 errored — inconclusive (comparison errors)

Scenario Net win Δ Pref Trials (W/T/L)
= Accurately classify NUnit tests with misleading method names +0.0% +0.0% 0/0/0
= Audit test distribution without modifying files +0.0% +0.0% 0/0/0
= Classify Go tests where no trait attribute mechanism exists +0.0% +0.0% 0/0/0
= Decline request to write new tests +0.0% +0.0% 0/0/0
= Tag MSTest tests and verify the project still builds +0.0% +0.0% 0/0/0
= Tag a partially-tagged MSTest suite without duplicating existing traits +0.0% +0.0% 0/0/0
= Tag an untagged MSTest test suite +0.0% +0.0% 0/0/0
= Tag an untagged NUnit test suite +0.0% +0.0% 0/0/0
= Tag an untagged xUnit test suite +0.0% +0.0% 0/0/0
⚠️ writing-mstest-tests — details

Reason: Net win +0.0% (0W/0T/0L over 0 trial(s), sign test p=1.000), mean preference +0.0%, 16 errored — inconclusive (comparison errors)

Scenario Net win Δ Pref Trials (W/T/L)
= Configure conditional execution, retry, and cleanup +0.0% +0.0% 0/0/0
= Configure test parallelization and MSTest.Sdk project +0.0% +0.0% 0/0/0
= Fix swapped Assert.AreEqual arguments +0.0% +0.0% 0/0/0
= Modernize legacy test patterns +0.0% +0.0% 0/0/0
= Replace ExpectedException with Assert.Throws +0.0% +0.0% 0/0/0
= Replace generic IsTrue checks for null, identity, emptiness, and absence +0.0% +0.0% 0/0/0
= Set up test lifecycle correctly +0.0% +0.0% 0/0/0
= Use DynamicData with ValueTuples over object arrays +0.0% +0.0% 0/0/0
= Use comparison assertions for boundary testing +0.0% +0.0% 0/0/0
= Use proper collection assertions +0.0% +0.0% 0/0/0
= Use proper type assertions instead of casts +0.0% +0.0% 0/0/0
= Use string assertions for format validation +0.0% +0.0% 0/0/0
= Write async tests with cancellation +0.0% +0.0% 0/0/0
= Write data-driven tests for a calculator +0.0% +0.0% 0/0/0
= Write tests with collection, null, and reference assertions +0.0% +0.0% 0/0/0
= Write unit tests for a service class +0.0% +0.0% 0/0/0

🔍 Full Results - additional metrics and failure investigation steps

▶ Sessions Visualisation -- interactive replay of all evaluation sessions
📊 Session Analytics (preview) -- aggregated metrics across evaluation sessions

@github-actions github-actions Bot added waiting-on-review PR state label and removed pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Aug 5, 2026
@github-actions

github-actions Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

✅ Evaluation passed for adbbc11. cc @AbhitejJohn @JanKrivanek — please review.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Pull requests that update a dependency file github_actions Pull requests that update GitHub Actions code waiting-on-review PR state label

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants