Inside one historical Mortal Kombat battle

Inside one historical Mortal Kombat battle A workflow diagram generated by Archify. 01 / One battle: ten example pairs, one weighted-rubric judgment EX / Failure branches in the preserved source Task + pair · ten example inputs · One battle: ten example pairs, one weighted-rubric judgment Task + pair ten example inputs Resolve A / B · per input: cache / call · One battle: ten example pairs, one weighted-rubric judgment Resolve A / B per input: cache / call Keep error · candidate failure · Failure branches in the preserved source Keep error candidate failure GPT-5.4 · whole-battle decision · One battle: ten example pairs, one weighted-rubric judgment GPT-5.4 whole-battle decision Judge error · request or parsing · Failure branches in the preserved source Judge error request or parsing Save record · battle + ten items · One battle: ten example pairs, one weighted-rubric judgment Save record battle + ten items ten pairs + rubric decision or DQ if failed include error if failed retain failure Legend Task selection Evaluator logic Failure handling Saved evidence Judge provider

What one decision means

  • • The request bundles the task, all ten source texts, paired outputs or errors, and the full 100-point rubric.
  • • GPT-5.4 returns a winner, reason and confidence; disqualification is also supported.
  • • No per-example judgments or numeric category scores are stored.

Read the evidence in context

  • • Preserved source flow; exact runtime revisions and particular cache hits remain unverified.
  • • The backfilled ranking is a separate endpoint, not an intermediate ladder replay.
  • • The companion showcase links two recorded wins to their examples, rubric and decisions.