Capture the work
Record the prompts, outputs, edits, evidence choices, timing, and model cost inside the challenge workspace.
How HILArena scores real work
Work inside one fixed, timed challenge. HILArena scores the final output and the decisions that shaped it, then shows you the evidence in plain language.
The scoring journey
The challenge conditions are fixed before you begin, and every score stays tied to the evidence that produced it.
Record the prompts, outputs, edits, evidence choices, timing, and model cost inside the challenge workspace.
Apply one fixed rubric to the final work and the decisions that produced it.
Identify where your judgment improved the AI-only reference and where it did not.
Codex and Fable review the same frozen evidence independently. Sunderam makes the final call when their judgments materially differ.
Publish the scores, evidence level, decision trail, and current use in plain language.
Independent reviews, human final call
Codex and Fable score the same finished work independently using one fixed rubric. Neither sees the other's judgment first. When they materially disagree, Sunderam reviews the evidence and records the final decision.
Evidence levels
Every result carries one clear status. Stronger uses open only for the exact evaluation versions that support them.
| Evidence level | What it means | What unlocks next | Availability |
|---|---|---|---|
| Experimental | Private feedback from a completed run with founder-supervised scoring. | Complete run records, traceable errors, consistent scoring, and feedback participants can use. | OPEN |
| Calibrated | Scores can be compared within the same challenge version and conditions. | Pre-set agreement, repeatability, stability, fairness, and appeal checks met with enough participants. | NOT ACTIVE |
| Validated | The result is approved for the exact credential or employer use shown beside it. | Use-specific outcome evidence, fairness review, ongoing monitoring, security, and regular re-checks. | NOT ACTIVE |
Built-in safeguards
You say yes or no to each of these separately: taking part, research use, public display, employer sharing, and any future AI-training use.
Your work, your identity, your exact rank, and your files stay private. They are shared only if you choose one of the sharing options the platform offers.
Each run record keeps the task, AI model, scoring rules, price, fixed test conditions, evaluator outputs, and evidence level together. That published record cannot be changed later.
Open the demo workspace — it runs on sample data — and follow one decision from start to score.