AimRank Ranking System · live at app.aimrank.io
Preferences in, calibrated rankings out, with the evidence attached
Anyone can compare two things. AimRank turns those comparisons into rankings that hold up: rating engines that carry their own uncertainty, agreement statistics across raters, confidence intervals, and a data card that says how the dataset was made.
This one is ours, not a client build. We run it, it has been public since July 2026, and the free tier is the whole product apart from the machine-readable exports. Agents get 21 read-only tools over MCP with no key and no account.
Free signups are open. No invite, no sales call.
What it does
Two halves that feed each other. One collects the preferences, the other uses them to grade an AI system.
Rank anything, pairwise
Text, image, audio and video. Every vote is one A versus B choice. Six rating engines compute the ratings: Glicko-2 by default, with Elo, Bradley-Terry, TrueSkill, OpenSkill and Elo-MMR selectable per ranking. Plackett-Luce and CrowdBT run as offline audit aggregators and never write ratings.
Preference data you can train on
The votes come out as an RLHF-ready preference dataset with per-sample provenance, ready for reward-model or preference-optimization training. The bundled export is the one thing the free tier does not reach.
Evals for your own AI
Suites, tasks and runs, graded by code checks, an LLM judge, a human reviewer scored by kappa, or the trajectory the agent took. Reliability comes back as pass^k with a Wilson interval, next to cost and latency, not as a single lucky run.
A record an auditor can read
Every ranking carries a data card, a rater-agreement panel and a hash-chained audit log, plus a methodology payload shaped like the Annex IV technical documentation. Uninstrumented is reported as unknown, never as zero.
What the free tier actually gives you
Signups are open and there is no paywall in the way. These are the numbers a free account actually meets.
A free account cannot mint an API key. That puts the keyed 56-tool MCP endpoint and the bundled RLHF export out of reach, because every keyed tool takes the key as an argument. Nothing else on this page is behind it, and the 21 read-only MCP tools need no account at all.
Agent-consumable
It speaks MCP, and 21 tools need no key
The public endpoint serves 21 read-only tools to any agent with no API key and no account: browse the catalogue, read any public ranking and its quality evidence, pull a matchup, get a win prediction and the explanation behind it. The keyed endpoint serves 56, which adds voting, creating rankings, running evals and export. Private rankings resolve as not found on the public endpoint, and a ranking is never more visible than the domain it sits in.
The tool count is 56 keyed and 21 public. Resources and Prompts are not tools and are not added to it.
The statistics, named
Every number the platform prints has a method behind it. These are the ones it runs.
Agreement between raters
Fleiss kappa, Scott's pi and Krippendorff alpha when there are many raters. Cohen's kappa is there too and is gated to exactly two raters, which is what it is defined for.
Intervals, not point estimates
Wilson 95% confidence intervals on rates, so three votes at 100% cannot pose as proven. Overlapping intervals are reported as no evidence of a difference, not as proof of equality.
Calibration, fairness and drift
Brier score and expected calibration error on the predicted win probability, four-fifths disparate impact across rater segments, and PSI and KS drift monitoring.
Complementarity
Whether the human and the model together beat either alone. Agreement statistics cannot answer that: a model that perfectly replicates the human scores a perfect kappa and adds nothing to the team.
pass^k
Agent reliability as the chance that k independent runs all pass, with a Wilson interval, reported beside cost and p95 latency so a reliability number cannot hide a regression.
A tamper-evident record
The audit log is hash-chained. Change a row and the chain breaks, and verification says so.
Where it runs
The platform runs on OVHcloud in Germany (DE1). The LLM judge runs on OVH AI Endpoints, model Qwen3.5-397B-A17B, so a judged comparison stays inside the EU. Sovereignty is the default here, not an add-on you buy.
Before you sign up
Read this first. It is cheaper for both of us.
- The public catalogue is three demo rankings They are named Demo, and their own descriptions say the votes were seeded to show the format, not cast by real people. The product works. The catalogue stays honest and thin until there is something worth arguing about in it.
- Paid plans are not open yet The pricing page on the app says so in its own words. The free tier is what you can have today, and it is deliberately generous. If you need the export or the keyed endpoint now, talk to us.
- The data card is aggregate It describes the dataset as a whole and it is not cryptographically signed. Per-sample provenance lives in the RLHF export, not on the card.
- Email and password only There is no Google or GitHub sign-in, and there is no live vote feed.
Go and use it
No invite and no sales call. Browse without an account, vote without an account, or sign up free and build your own ranking.
Need the bundled export or the keyed endpoint now? Talk to us.