Show HN: Human Benchmark – Compare your reasoning skills against AI models
Human Benchmark is an interactive platform that evaluates your performance by answering questions used to measure AI reasoning abilities. It adapts difficulty based on your ability and times responses. Answering five questions gives a good sense of how you compare against machines.
Human Benchmark — how do you score against the machines?
Answer questions used to measure AI reasoning abilities
Difficulty adapts according to your ability
Responses are timed
Five questions will give a good sense of how you compare (and your potential earnings)
Time remaining—
Question 1—
—
That's your measure
The Leaderboard
— models
Shaded bars are confidence ranges.
— items
Time per question 1m 00s
Run out of time and the question is marked wrong.
Changes apply to the next question.
Question Sources
CommonsenseQA — everyday common sense that most people find easy.
GSM8K — grade-school arithmetic word problems.
AGIEval (LSAT logical reasoning) — real law-school admission test items: read a short argument, spot the flaw or the assumption.
AGIEval (AQuA-RAT) — GMAT and GRE style quantitative word problems.
BIG-Bench Hard — logical deduction, object counting, date arithmetic, tracking things as they move, causal judgement, and truth-teller puzzles.
Methodology
The app uses a Rasch model to estimate a percentage score from a relatively small number of items. Information
Reading the range
The shaded bar is a 95% confidence range. It starts enormous and narrows with every answer. If two bars overlap heavily, the difference between them is within the margin of error.
Disclaimers
This is a proof of concept and illustration of how LLM benchmarking works in a human context, but it is not a serious test of reasoning or intelligence and has limited psychometric validty.