The Unjournal research prioritization prototype
Back to the paper dashboard
Methods note and live pilot

How stable are the AI priority scores?

A useful score should not swing sharply because the same model was run again, a harmless instruction was rephrased, or a comparable model was used. We are adding explicit tests for those forms of fragility.

Stability is not accuracy. A model can repeat the same mistaken judgment every time. Prompt stability is therefore an early reproducibility check; comparison with existing human judgments remains a separate, descriptive test.

The idea in three checks

The approach adapts inter-rater reliability: repeated runs, prompt variants, or models play the role of different raters scoring the same frozen set of papers.

Check 1

Measure ordinary rerun variation

Run each wording twice with the paper, rubric, model, settings, and output schema fixed. Differences within a wording estimate ordinary run-to-run variation.

Check 2

Separate wording sensitivity

Average the two runs within each reviewed, equivalent wording, then compare those wording means. This reduces ordinary rerun noise in the wording comparison.

Separate validation

Compare with people and models

Compare each model's stability profile and then compare its ratings with existing team judgments. Only the second comparison speaks to accuracy or calibration.

Variation within each wording
Difference between wording means
Validation against human judgments

Live pilot status

The page publishes aggregate diagnostics only. It does not expose paper-level reruns, rationales, or paper-level human ratings.

What is being tested: every stability number below describes AI model scores. “Previously human-rated papers” names a paper cohort selected because earlier Unjournal ratings are available; those ratings do not enter the stability calculations. Human and AI scores are compared separately in the amber reference panel.

Balanced crossed design: two wordings × two executions

Each paper receives four AI scores. Only the final closing instruction changes; all other scoring inputs remain fixed.

Closing instructionExecution 1Execution 2
Baseline wordingB1B2
Cautious wordingC1C2
  • Ordinary rerun variation: compare B1 with B2 and C1 with C2, separately.
  • Wording sensitivity: compare the average of B1/B2 with the average of C1/C2.
  • Held-out check: a third baseline execution is kept outside the balanced headline analysis and used only as a robustness check.

Two repeats reduce the confounding between wording and ordinary model variation, but they do not estimate the full range of possible prompts or executions. Treating the two instructions as substantively equivalent is a reviewed design judgment, not a fact established by this pilot.

LoadingReading pilot data...

What the metrics measure

Agreement on numeric scores

Interval Krippendorff alpha for the overall 0-100 rating and each 0-10 component. The paper treats about 0.8 as a useful, context-dependent reference point, not a universal cutoff.

Operationally important movement

The mean, median, and maximum within-paper score range, plus the number of papers that cross 25, 50, 65, or 75 across runs. A modest wobble can matter if it changes a workflow decision.

Ranking and action stability

Mean pairwise rank correlation and nominal agreement on recommended action. A high alpha can coexist with a few important shortlist reversals, so both views matter.

Why not display the model's own confidence as the answer?

A self-reported confidence number is another model output. It may be useful, but it does not show how the pipeline actually behaves under reruns or prompt changes. Observed stability measures that behavior directly.

Metric glossary: what the statistics mean
Prompt Stability Score (PSS)
A family of observed-agreement tests. It asks whether AI ratings change when we rerun the same prompt or make a meaning-preserving wording change. It does not test whether the ratings are correct.
Intra-prompt PSS
Agreement among repeated AI runs with wording held fixed. Here it is estimated separately for the baseline and cautious wordings, rather than assuming one wording represents all run variation.
Inter-prompt PSS
Agreement between reviewed, equivalent wordings. Here each wording score is the mean of two executions, so the comparison is less affected by ordinary run-to-run variation. The long scoring rubric stays fixed.
Crossed design
Every paper is scored under every wording × execution combination. This lets us compare the observed wording shift with movement that occurs even when wording is unchanged.
Mean absolute wording effect
For each paper, take the absolute difference between its cautious-wording mean and baseline-wording mean, then average across papers. It measures the size, not the direction, of this wording contrast.
Mean signed wording effect
The average cautious-minus-baseline difference. A negative value means the cautious wording lowered scores on average; an interval spanning zero does not show a clear directional shift.
Krippendorff alpha
An agreement statistic adjusted for disagreement expected from the pooled ratings. 1 means exact agreement; 0 means observed disagreement equals expected disagreement; negative values indicate more disagreement than expected. About 0.8 is a useful reference here, not a universal pass/fail rule.
95% paper-bootstrap interval
An uncertainty interval made by repeatedly resampling whole papers and recomputing alpha. Keeping each paper's reruns together respects the repeated-rating structure. Very small cohorts produce wide, sometimes negative intervals.
Within-paper range
The highest AI priority score minus the lowest for the same paper under the selected test. We report the mean and median across papers, plus the maximum fragile case.
Rank correlation
Mean pairwise Spearman correlation between AI run rankings. 1 preserves the paper order exactly; 0 indicates no monotonic relationship; negative values reverse the order.
Recommended-action alpha
Chance-adjusted agreement on categorical AI recommendations, such as shortlist or do not prioritize. This uses nominal rather than numeric alpha because the action labels are categories.
Same broad action tier
The share that stayed in the same broad action tier (below 25, 25–49, 50–74, or 75+). Ordinary-rerun summaries use paper–wording pairs; wording-sensitivity summaries use papers. It is an operational summary, not a statistical reliability coefficient.
Boundary flip
The count and share of papers with AI scores on both sides of a dashboard boundary (25, 50, 65, or 75) across reruns. Even modest score movement can matter near a workflow threshold.

How this is adapted to Unjournal prioritization

  1. Freeze the inputs. Use the same public title, abstract, authors, source, publication status, field hint, and cause-area hint for every condition.
  2. Use two deliberate cohorts. One contains newer AI human-impact/governance papers. The comparison cohort contains four historical AI/social-governance papers and two forecasting/x-risk papers with prior team ratings.
  3. Cross wording with execution. Run each of two reviewed closing instructions twice. The long prioritization rubric, aggregate human-calibration guidance, and JSON schema remain unchanged.
  4. Use one model first. Establish whether the current production setup is internally repeatable before spending effort comparing providers or model families.
  5. Do not mix fallbacks. A failed run stays failed; it is not silently replaced with a different model. Otherwise a supposed prompt test would partly become a model test.
  6. Publish aggregates, inspect cases privately. The public page shows reliability and boundary diagnostics. The team can inspect the largest disagreements to improve the prompt, metadata, or rubric.

This first pilot is intentionally narrower than paraphrasing the entire prompt. The stable core contains policy choices and calibration anchors; changing those would test a different construct, not harmless wording.

Exact closing instructions used in this pilot

These are the two closing instructions used in the balanced crossed test. Each was executed twice. The paper text, long Unjournal scoring rubric, calibration guidance, component definitions, and JSON output schema stayed fixed.

Baseline (baseline) Be calibrated, conservative, and honest about scope fit.
Cautious plain language (cautious_plain) Keep the assessment well calibrated. Avoid optimistic score inflation and describe the paper's fit with Unjournal's scope candidly.

How results should change the project

FindingLikely diagnosisPractical response
Low repeatability within a wordingRun-level nondeterminism, ambiguous papers, underspecified criteria, or inadequate model capability.Inspect the largest same-wording ranges, tighten edge-case rules, and add repetitions before drawing prompt conclusions.
Runs disagree about existing public scrutinyThe scorer is reconstructing a consequential factual input differently on each run.Supply a deterministic scrutiny assessment or verified evidence bundle, and require the scorer to mark missing evidence as unclear rather than infer it afresh.
Stable within wordings but different wording meansThe closing-instruction contrast is changing the judgment beyond ordinary rerun movement.Review the wording-sensitive cases and revise or standardize the fragile instruction before relying on sharp thresholds.
High alpha but many 65-point flipsOverall agreement looks good, but operational decisions near a boundary remain fragile.Show a stability flag near the threshold, request a human second look, or avoid treating 65 as a mechanically sharp boundary.
Stable but poorly matched to humansThe system is reproducibly wrong or systematically miscalibrated.Change calibration anchors or rubric interpretation; prompt stability alone offers no reassurance.
Models disagreeCould reflect capability, training, or interpretation differences.Compare each model's own stability and its error against blinded human anchors; do not choose from agreement alone.

Sources and implementation

Implementation note: the dashboard adaptation uses paper-level bootstrap resampling so the repeated ratings for one paper stay together. This is a deliberate variation from the package implementation described in the paper's appendix, which resamples long-format annotation records.