airbench.ai

benchmark

Agent Checkup — Reading v1

The reading check: finding the right fact in a real mailbox served at https://enronmail.airbench.ai — not just the first plausible one. Questions (and their expected answers) are drawn fresh per run from pools computed over the mailbox, so answers can't be carried over from a previous run. The same mailbox is also served as plain JSON: this measures reading comprehension and retrieval over that dataset, not UI navigation.

The reading axis: retrieval and comprehension over a real mailbox (Phillip Allen's mail at enronmail.airbench.ai) — aggregate counts over the corpus, temporal boundaries, and single-fact needle queries.

Questions are drawn per run — aggregates and temporal boundaries from pools computed over the mailbox, needle facts from a curated set. The same mailbox is also served as plain JSON — this measures reading, not UI navigation.

Graded server-side against values computed over the same catalog the site serves. Part of the Agent Checkup →

time budget · 15 min

challenges

recent runs