AI News HubLIVE
In-site rewrite1 min read

Show HN: Human Benchmark – Compare your reasoning skills against AI models

Human Benchmark is an interactive platform that evaluates your performance by answering questions used to measure AI reasoning abilities. It adapts difficulty based on your ability and times responses. Answering five questions gives a good sense of how you compare against machines.

SourceHacker News AIAuthor: twoslide

Human Benchmark — how do you score against the machines?

Answer questions used to measure AI reasoning abilities

Difficulty adapts according to your ability

Responses are timed

Five questions will give a good sense of how you compare (and your potential earnings)

Time remaining—

Question 1—

That's your measure

The Leaderboard

— models

Shaded bars are confidence ranges.

— items

Time per question 1m 00s

Run out of time and the question is marked wrong.

Changes apply to the next question.

Question Sources

CommonsenseQA — everyday common sense that most people find easy.

GSM8K — grade-school arithmetic word problems.

AGIEval (LSAT logical reasoning) — real law-school admission test items: read a short argument, spot the flaw or the assumption.

AGIEval (AQuA-RAT) — GMAT and GRE style quantitative word problems.

BIG-Bench Hard — logical deduction, object counting, date arithmetic, tracking things as they move, causal judgement, and truth-teller puzzles.

Methodology

The app uses a Rasch model to estimate a percentage score from a relatively small number of items. Information

Reading the range

The shaded bar is a 95% confidence range. It starts enormous and narrows with every answer. If two bars overlap heavily, the difference between them is within the margin of error.

Disclaimers

This is a proof of concept and illustration of how LLM benchmarking works in a human context, but it is not a serious test of reasoning or intelligence and has limited psychometric validty.