AimRank Ranking System · live at app.aimrank.io

Preferences in, calibrated rankings out, with the evidence attached

Anyone can compare two things. AimRank turns those comparisons into rankings that hold up: rating engines that carry their own uncertainty, agreement statistics across raters, confidence intervals, and a data card that says how the dataset was made.

This one is ours, not a client build. We run it, it has been public since July 2026, and the free tier is the whole product apart from the machine-readable exports. Agents get 21 read-only tools over MCP with no key and no account.

Free signups are open. No invite, no sales call.

6 rating engines, plus 2 offline audit aggregators
21 MCP tools with no API key
56 MCP tools on the keyed endpoint
€0 free tier, signups open

What it does

Two halves that feed each other. One collects the preferences, the other uses them to grade an AI system.

Rank anything, pairwise

Text, image, audio and video. Every vote is one A versus B choice. Six rating engines compute the ratings: Glicko-2 by default, with Elo, Bradley-Terry, TrueSkill, OpenSkill and Elo-MMR selectable per ranking. Plackett-Luce and CrowdBT run as offline audit aggregators and never write ratings.

Preference data you can train on

The votes come out as an RLHF-ready preference dataset with per-sample provenance, ready for reward-model or preference-optimization training. The bundled export is the one thing the free tier does not reach.

Evals for your own AI

Suites, tasks and runs, graded by code checks, an LLM judge, a human reviewer scored by kappa, or the trajectory the agent took. Reliability comes back as pass^k with a Wilson interval, next to cost and latency, not as a single lucky run.

A record an auditor can read

Every ranking carries a data card, a rater-agreement panel and a hash-chained audit log, plus a methodology payload shaped like the Annex IV technical documentation. Uninstrumented is reported as unknown, never as zero.

What the free tier actually gives you

Signups are open and there is no paywall in the way. These are the numbers a free account actually meets.

Rankings 10 active
Entities per ranking 100
Votes 1,000 per month
Eval suites 25
Eval runs 200 per month
LLM-judge calls 50 per month
Media types Text, image, audio and video. None of them is gated.
Rating engines All six, selectable per ranking.
The statistics Data card, rater agreement, confidence intervals, calibration, fairness, drift, complementarity and the hash-chained audit log. All of it.
The one wall

A free account cannot mint an API key. That puts the keyed 56-tool MCP endpoint and the bundled RLHF export out of reach, because every keyed tool takes the key as an argument. Nothing else on this page is behind it, and the 21 read-only MCP tools need no account at all.

Agent-consumable

It speaks MCP, and 21 tools need no key

The public endpoint serves 21 read-only tools to any agent with no API key and no account: browse the catalogue, read any public ranking and its quality evidence, pull a matchup, get a win prediction and the explanation behind it. The keyed endpoint serves 56, which adds voting, creating rankings, running evals and export. Private rankings resolve as not found on the public endpoint, and a ranking is never more visible than the domain it sits in.

Public endpoint https://app.aimrank.io/mcp/public · 21 read-only tools · no key, no account
Keyed endpoint https://app.aimrank.io/mcp/mcp · 56 tools · the key is a tool argument, not an HTTP header
Registry Listed on the official MCP registry as io.aimrank/aimrank
Also served 2 Resources, 3 Resource templates and 5 Prompts. Separate surfaces, never counted as tools.

The tool count is 56 keyed and 21 public. Resources and Prompts are not tools and are not added to it.

The statistics, named

Every number the platform prints has a method behind it. These are the ones it runs.

Agreement between raters

Fleiss kappa, Scott's pi and Krippendorff alpha when there are many raters. Cohen's kappa is there too and is gated to exactly two raters, which is what it is defined for.

Intervals, not point estimates

Wilson 95% confidence intervals on rates, so three votes at 100% cannot pose as proven. Overlapping intervals are reported as no evidence of a difference, not as proof of equality.

Calibration, fairness and drift

Brier score and expected calibration error on the predicted win probability, four-fifths disparate impact across rater segments, and PSI and KS drift monitoring.

Complementarity

Whether the human and the model together beat either alone. Agreement statistics cannot answer that: a model that perfectly replicates the human scores a perfect kappa and adds nothing to the team.

pass^k

Agent reliability as the chance that k independent runs all pass, with a Wilson interval, reported beside cost and p95 latency so a reliability number cannot hide a regression.

A tamper-evident record

The audit log is hash-chained. Change a row and the chain breaks, and verification says so.

Where it runs

The platform runs on OVHcloud in Germany (DE1). The LLM judge runs on OVH AI Endpoints, model Qwen3.5-397B-A17B, so a judged comparison stays inside the EU. Sovereignty is the default here, not an add-on you buy.

Before you sign up

Read this first. It is cheaper for both of us.

  • The public catalogue is three demo rankings They are named Demo, and their own descriptions say the votes were seeded to show the format, not cast by real people. The product works. The catalogue stays honest and thin until there is something worth arguing about in it.
  • Paid plans are not open yet The pricing page on the app says so in its own words. The free tier is what you can have today, and it is deliberately generous. If you need the export or the keyed endpoint now, talk to us.
  • The data card is aggregate It describes the dataset as a whole and it is not cryptographically signed. Per-sample provenance lives in the RLHF export, not on the card.
  • Email and password only There is no Google or GitHub sign-in, and there is no live vote feed.

Go and use it

No invite and no sales call. Browse without an account, vote without an account, or sign up free and build your own ranking.

Need the bundled export or the keyed endpoint now? Talk to us.