Measure ordinary rerun variation
Run each wording twice with the paper, rubric, model, settings, and output schema fixed. Differences within a wording estimate ordinary run-to-run variation.
A useful score should not swing sharply because the same model was run again, a harmless instruction was rephrased, or a comparable model was used. We are adding explicit tests for those forms of fragility.
The approach adapts inter-rater reliability: repeated runs, prompt variants, or models play the role of different raters scoring the same frozen set of papers.
Run each wording twice with the paper, rubric, model, settings, and output schema fixed. Differences within a wording estimate ordinary run-to-run variation.
Average the two runs within each reviewed, equivalent wording, then compare those wording means. This reduces ordinary rerun noise in the wording comparison.
Compare each model's stability profile and then compare its ratings with existing team judgments. Only the second comparison speaks to accuracy or calibration.
The page publishes aggregate diagnostics only. It does not expose paper-level reruns, rationales, or paper-level human ratings.
Each paper receives four AI scores. Only the final closing instruction changes; all other scoring inputs remain fixed.
| Closing instruction | Execution 1 | Execution 2 |
|---|---|---|
| Baseline wording | B1 | B2 |
| Cautious wording | C1 | C2 |
Two repeats reduce the confounding between wording and ordinary model variation, but they do not estimate the full range of possible prompts or executions. Treating the two instructions as substantively equivalent is a reviewed design judgment, not a fact established by this pilot.
Interval Krippendorff alpha for the overall 0-100 rating and each 0-10 component. The paper treats about 0.8 as a useful, context-dependent reference point, not a universal cutoff.
The mean, median, and maximum within-paper score range, plus the number of papers that cross 25, 50, 65, or 75 across runs. A modest wobble can matter if it changes a workflow decision.
Mean pairwise rank correlation and nominal agreement on recommended action. A high alpha can coexist with a few important shortlist reversals, so both views matter.
A self-reported confidence number is another model output. It may be useful, but it does not show how the pipeline actually behaves under reruns or prompt changes. Observed stability measures that behavior directly.
This first pilot is intentionally narrower than paraphrasing the entire prompt. The stable core contains policy choices and calibration anchors; changing those would test a different construct, not harmless wording.
These are the two closing instructions used in the balanced crossed test. Each was executed twice. The paper text, long Unjournal scoring rubric, calibration guidance, component definitions, and JSON output schema stayed fixed.
baseline)
Be calibrated, conservative, and honest about scope fit.
cautious_plain)
Keep the assessment well calibrated. Avoid optimistic score inflation and describe the paper's fit with Unjournal's scope candidly.
| Finding | Likely diagnosis | Practical response |
|---|---|---|
| Low repeatability within a wording | Run-level nondeterminism, ambiguous papers, underspecified criteria, or inadequate model capability. | Inspect the largest same-wording ranges, tighten edge-case rules, and add repetitions before drawing prompt conclusions. |
| Runs disagree about existing public scrutiny | The scorer is reconstructing a consequential factual input differently on each run. | Supply a deterministic scrutiny assessment or verified evidence bundle, and require the scorer to mark missing evidence as unclear rather than infer it afresh. |
| Stable within wordings but different wording means | The closing-instruction contrast is changing the judgment beyond ordinary rerun movement. | Review the wording-sensitive cases and revise or standardize the fragile instruction before relying on sharp thresholds. |
| High alpha but many 65-point flips | Overall agreement looks good, but operational decisions near a boundary remain fragile. | Show a stability flag near the threshold, request a human second look, or avoid treating 65 as a mechanically sharp boundary. |
| Stable but poorly matched to humans | The system is reproducibly wrong or systematically miscalibrated. | Change calibration anchors or rubric interpretation; prompt stability alone offers no reassurance. |
| Models disagree | Could reflect capability, training, or interpretation differences. | Compare each model's own stability and its error against blinded human anchors; do not choose from agreement alone. |
Implementation note: the dashboard adaptation uses paper-level bootstrap resampling so the repeated ratings for one paper stay together. This is a deliberate variation from the package implementation described in the paper's appendix, which resamples long-format annotation records.