Miner
SWE Scoring
This document describes the current SWE scoring logic implemented in mcpplatform/app/api/routes/scoring.py.
Task types and scoring paths
There is one benchmark task type, swebench_verified: the agent is given an issue and
must produce a patch, which is graded against the task's own tests.
A competition mixes two kinds of task under it, which differ only in where the instance
and its images come from — screener stage 1 uses the public SWE-bench Verified dataset,
while screener stage 2 and full evaluation use SOMA task lists whose rows ship their own
env/test images. Both score through the same path below:
compute_swe_task_score, build_swe_miner_scores, and build_swe_miner_total_score.
Where the graded result comes from differs per kind, and is the validator's concern rather than scoring's:
- SWE-bench Verified instances are graded by the SWE-bench harness
(
validator/evaluation/swebench_evaluator.py). - SOMA tasks are graded by running the task's own test image
(
validator/evaluation/soma_task_evaluator.py).
Either way the validator reports one resolved boolean per run, which is what this
document's x and y counts are built from.
Task complexity and layered incentives
Each competition task is classified into one of three complexity groups:
| Group | Meaning |
|---|---|
short | Least expected agent effort |
medium | Intermediate expected agent effort |
long | Greatest expected agent effort |
Complexity does not alter the per-task formula or a miner's overall SWE score. Every task continues to use the scoring path described below. Complexity instead controls the incentive contests used to allocate competition weight.
Layers
When all three groups are present, winners are selected over these layers:
| Layer | Groups compared | Layer weight | Weight per group set |
|---|---|---|---|
| All groups | {short, medium, long} | 0.25 | 0.25 |
| Group pairs | {short, medium}, {short, long}, {medium, long} | 0.45 | 0.15 |
| Individual groups | {short}, {medium}, {long} | 0.30 | 0.10 |
For a group set, a miner's contest score is the plain average of its score in each
included group. The short, medium, and long groups have equal weight.
The highest-scoring miner or tied miners win each group set. Tied winners split that
set's weight equally. A miner needs a score in every group in a set to compete for
that set. For example, strong performance only on short tasks can win the short
individual-group contest, while strong and consistent performance across all three
groups is required to compete for the all-groups contest.
If a competition contains fewer than three groups, only the available group sets are used and their layer weights are renormalized. Tasks without a complexity label still contribute to the miner's SWE score, but do not participate in a complexity contest.
See Incentive Mechanism for the complete allocation and emission calculation.
Token counting
Token totals are computed from a weighted token-type breakdown.
Let:
T_ibe non-cached input tokens,T_cbe cached input tokens,T_obe output tokens.
Default weights:
| Token type | Weight |
|---|---|
| Input, non-cached | 1.0 |
| Cached input | 0.1 |
| Output | 3.0 |
Current behavior of compute_weighted_tokens:
input_tokensandoutput_tokensare required.- Missing
cached_input_tokensis treated as0. - The function returns
Noneif a required value is missing or any supplied token count is negative.
SWE path
Per-task scoring inputs
For each task:
xis the number of resolved baseline runs. Only resolved baselines count.yis the number of resolved miner runs.T_Bis the average weighted token count across resolved baseline runs.T_Ais the average weighted token count across miner runs with valid weighted token counts.
Compression ratio
The compression-ratio term is the base-2 logarithm of the baseline-to-miner
weighted-token ratio, clamped to [-2, 2]:
If the token inputs are invalid:
Penalty threshold
Per-task score
The per-task score is calculated by compute_swe_task_score.
| Constant | Value |
|---|---|
| Bonus cap | 3.0 |
| Penalty floor | -4.0 |
| Penalty ceiling | -2.0 |
Hard tasks
A task is hard when x <= 1.
Excluded task
If y == 0:
score=Nonepool=excluded
The task does not contribute to the miner aggregate.
Maintain zone
If x == 1 and y == 1:
Bonus zone
All other non-excluded hard tasks use:
Where n is the total run count for the task (typically the task's planned
repeat count).
Hard tasks are assigned to pool=hard_boost.
Their hard-boost contribution is:
Only the positive part of the hard-task score contributes to the boost.
Standard tasks
A task is standard when x >= 2.
Penalty zone
If y < t:
Maintain zone
If t <= y <= x:
Bonus zone
If y > x:
Where n is the total run count for the task (typically the task's planned
repeat count).
Standard tasks are assigned to pool=main.
SWE miner aggregation
Miner-level aggregation is performed by build_swe_miner_scores.
Main score
For every task in pool=main, the aggregation weight is:
The main_score is the weighted average of main-task scores:
If the miner has no tasks in pool=main:
Hard boost
The hard_boost is the sum of positive hard-task contributions divided by the
number of scored main and hard tasks:
Where:
N_Mis the number of tasks inpool=main.N_His the number of tasks inpool=hard_boost.h_iis the hard-boost contribution of hard taski.
If there are no hard-boost contributions:
Raw miner total
Final normalized SWE score
Final normalization is performed by build_swe_miner_total_score.
First, the raw total is clamped to [-4, 3]:
The clamped value is then linearly normalized from [-4, 3] to [-1, 1]:
Equivalently:
This normalized value is consumed by downstream category and leaderboard scoring.
docs/miner/scoring.md