How HILArena scores real work

See exactly how
your result was earned.

Work inside one fixed, timed challenge. HILArena scores the final output and the decisions that shaped it, then shows you the evidence in plain language.

Work + decisionsHuman vs referenceReview you can challenge

The scoring journey

From your attempt to a result you can inspect.

The challenge conditions are fixed before you begin, and every score stays tied to the evidence that produced it.

01

Capture the work

Record the prompts, outputs, edits, evidence choices, timing, and model cost inside the challenge workspace.

02

Score what mattered

Apply one fixed rubric to the final work and the decisions that produced it.

03

Show the difference

Identify where your judgment improved the AI-only reference and where it did not.

04

Resolve disagreement

Codex and Fable review the same frozen evidence independently. Sunderam makes the final call when their judgments materially differ.

05

Explain the result

Publish the scores, evidence level, decision trail, and current use in plain language.

Independent reviews, human final call

Important disagreements get reviewed.

Codex and Fable score the same finished work independently using one fixed rubric. Neither sees the other's judgment first. When they materially disagree, Sunderam reviews the evidence and records the final decision.

Participant run
CodexIndependent review A
FableIndependent review B
SunderamFinal decision

Evidence levels

Know what each result is ready for.

Every result carries one clear status. Stronger uses open only for the exact evaluation versions that support them.

Evidence levelWhat it meansWhat unlocks nextAvailability
ExperimentalPrivate feedback from a completed run with founder-supervised scoring.Complete run records, traceable errors, consistent scoring, and feedback participants can use.OPEN
CalibratedScores can be compared within the same challenge version and conditions.Pre-set agreement, repeatability, stability, fairness, and appeal checks met with enough participants.NOT ACTIVE
ValidatedThe result is approved for the exact credential or employer use shown beside it.Use-specific outcome evidence, fairness review, ongoing monitoring, security, and regular re-checks.NOT ACTIVE

Built-in safeguards

Trust is built into every run.

Consent is separated

You say yes or no to each of these separately: taking part, research use, public display, employer sharing, and any future AI-training use.

Private by default

Your work, your identity, your exact rank, and your files stay private. They are shared only if you choose one of the sharing options the platform offers.

Version everything

Each run record keeps the task, AI model, scoring rules, price, fixed test conditions, evaluator outputs, and evidence level together. That published record cannot be changed later.

Want to see how it works, not just read about it?

Open the demo workspace — it runs on sample data — and follow one decision from start to score.

Try the demo