<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom"><title>FronteraEval updates</title><id>https://fronteraeval.org/</id><link href="https://fronteraeval.org/feed.xml" rel="self"/><link href="https://fronteraeval.org/"/><updated>2026-08-30T08:25:18.742Z</updated><subtitle>New and updated frontier AI evaluation records.</subtitle><entry><title>AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions</title><id>https://fronteraeval.org/evaluations/inspect-abstention-bench/</id><link href="https://fronteraeval.org/evaluations/inspect-abstention-bench/"/><updated>2026-08-30T00:00:00Z</updated><summary>Evaluating abstention across 20 diverse datasets, including questions with unknown answers, underspecification, false premises, subjective interpretations, and outdated information.</summary></entry><entry><title>AgentBench: Evaluate LLMs as Agents</title><id>https://fronteraeval.org/evaluations/inspect-agent-bench-os/</id><link href="https://fronteraeval.org/evaluations/inspect-agent-bench-os/"/><updated>2026-08-30T00:00:00Z</updated><summary>A benchmark designed to evaluate LLMs as Agents</summary></entry><entry><title>Agent Threat Bench Autonomy Hijack</title><id>https://fronteraeval.org/evaluations/inspect-agent-threat-bench-autonomy-hijack/</id><link href="https://fronteraeval.org/evaluations/inspect-agent-threat-bench-autonomy-hijack/"/><updated>2026-08-30T00:00:00Z</updated><summary>Evaluates LLM agents against the OWASP Top 10 for Agentic Applications (2026), measuring both task utility and security resilience across memory poisoning, autonomy hijacking, and data exfiltration scenarios.</summary></entry><entry><title>Agent Threat Bench Data Exfil</title><id>https://fronteraeval.org/evaluations/inspect-agent-threat-bench-data-exfil/</id><link href="https://fronteraeval.org/evaluations/inspect-agent-threat-bench-data-exfil/"/><updated>2026-08-30T00:00:00Z</updated><summary>Evaluates LLM agents against the OWASP Top 10 for Agentic Applications (2026), measuring both task utility and security resilience across memory poisoning, autonomy hijacking, and data exfiltration scenarios.</summary></entry><entry><title>Agent Threat Bench Memory Poison</title><id>https://fronteraeval.org/evaluations/inspect-agent-threat-bench-memory-poison/</id><link href="https://fronteraeval.org/evaluations/inspect-agent-threat-bench-memory-poison/"/><updated>2026-08-30T00:00:00Z</updated><summary>Evaluates LLM agents against the OWASP Top 10 for Agentic Applications (2026), measuring both task utility and security resilience across memory poisoning, autonomy hijacking, and data exfiltration scenarios.</summary></entry><entry><title>AgentBoard</title><id>https://fronteraeval.org/evaluations/canonical-agentboard/</id><link href="https://fronteraeval.org/evaluations/canonical-agentboard/"/><updated>2026-08-30T00:00:00Z</updated><summary>Multi-environment benchmark and analysis toolkit for language-model agents.</summary></entry><entry><title>AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents</title><id>https://fronteraeval.org/evaluations/inspect-agentdojo/</id><link href="https://fronteraeval.org/evaluations/inspect-agentdojo/"/><updated>2026-08-30T00:00:00Z</updated><summary>Assesses whether AI agents can be hijacked by malicious third parties using prompt injections in simple environments such as a workspace or travel booking app.</summary></entry><entry><title>AgentHarm</title><id>https://fronteraeval.org/evaluations/canonical-agentharm/</id><link href="https://fronteraeval.org/evaluations/canonical-agentharm/"/><updated>2026-08-30T00:00:00Z</updated><summary>Evaluates whether language-model agents can execute harmful multi-step tasks.</summary></entry><entry><title>Agentharm Benign</title><id>https://fronteraeval.org/evaluations/inspect-agentharm-benign/</id><link href="https://fronteraeval.org/evaluations/inspect-agentharm-benign/"/><updated>2026-08-30T00:00:00Z</updated><summary>Assesses whether AI agents might engage in harmful activities by testing their responses to malicious prompts in areas like cybercrime, harassment, and fraud, aiming to ensure safe behavior.</summary></entry><entry><title>Agie Aqua Rat</title><id>https://fronteraeval.org/evaluations/inspect-agie-aqua-rat/</id><link href="https://fronteraeval.org/evaluations/inspect-agie-aqua-rat/"/><updated>2026-08-30T00:00:00Z</updated><summary>AGIEval is a human-centric benchmark specifically designed to evaluate the general abilities of foundation models in tasks pertinent to human cognition and problem-solving.</summary></entry><entry><title>Agie Logiqa En</title><id>https://fronteraeval.org/evaluations/inspect-agie-logiqa-en/</id><link href="https://fronteraeval.org/evaluations/inspect-agie-logiqa-en/"/><updated>2026-08-30T00:00:00Z</updated><summary>AGIEval is a human-centric benchmark specifically designed to evaluate the general abilities of foundation models in tasks pertinent to human cognition and problem-solving.</summary></entry><entry><title>Agie Lsat Ar</title><id>https://fronteraeval.org/evaluations/inspect-agie-lsat-ar/</id><link href="https://fronteraeval.org/evaluations/inspect-agie-lsat-ar/"/><updated>2026-08-30T00:00:00Z</updated><summary>AGIEval is a human-centric benchmark specifically designed to evaluate the general abilities of foundation models in tasks pertinent to human cognition and problem-solving.</summary></entry><entry><title>Agie Lsat Lr</title><id>https://fronteraeval.org/evaluations/inspect-agie-lsat-lr/</id><link href="https://fronteraeval.org/evaluations/inspect-agie-lsat-lr/"/><updated>2026-08-30T00:00:00Z</updated><summary>AGIEval is a human-centric benchmark specifically designed to evaluate the general abilities of foundation models in tasks pertinent to human cognition and problem-solving.</summary></entry><entry><title>Agie Lsat Rc</title><id>https://fronteraeval.org/evaluations/inspect-agie-lsat-rc/</id><link href="https://fronteraeval.org/evaluations/inspect-agie-lsat-rc/"/><updated>2026-08-30T00:00:00Z</updated><summary>AGIEval is a human-centric benchmark specifically designed to evaluate the general abilities of foundation models in tasks pertinent to human cognition and problem-solving.</summary></entry><entry><title>Agie Math</title><id>https://fronteraeval.org/evaluations/inspect-agie-math/</id><link href="https://fronteraeval.org/evaluations/inspect-agie-math/"/><updated>2026-08-30T00:00:00Z</updated><summary>AGIEval is a human-centric benchmark specifically designed to evaluate the general abilities of foundation models in tasks pertinent to human cognition and problem-solving.</summary></entry><entry><title>Agie Sat En</title><id>https://fronteraeval.org/evaluations/inspect-agie-sat-en/</id><link href="https://fronteraeval.org/evaluations/inspect-agie-sat-en/"/><updated>2026-08-30T00:00:00Z</updated><summary>AGIEval is a human-centric benchmark specifically designed to evaluate the general abilities of foundation models in tasks pertinent to human cognition and problem-solving.</summary></entry><entry><title>Agie Sat En Without Passage</title><id>https://fronteraeval.org/evaluations/inspect-agie-sat-en-without-passage/</id><link href="https://fronteraeval.org/evaluations/inspect-agie-sat-en-without-passage/"/><updated>2026-08-30T00:00:00Z</updated><summary>AGIEval is a human-centric benchmark specifically designed to evaluate the general abilities of foundation models in tasks pertinent to human cognition and problem-solving.</summary></entry><entry><title>Agie Sat Math</title><id>https://fronteraeval.org/evaluations/inspect-agie-sat-math/</id><link href="https://fronteraeval.org/evaluations/inspect-agie-sat-math/"/><updated>2026-08-30T00:00:00Z</updated><summary>AGIEval is a human-centric benchmark specifically designed to evaluate the general abilities of foundation models in tasks pertinent to human cognition and problem-solving.</summary></entry><entry><title>AHB</title><id>https://fronteraeval.org/evaluations/register-ahb/</id><link href="https://fronteraeval.org/evaluations/register-ahb/"/><updated>2026-08-30T00:00:00Z</updated><summary>A text-only safety benchmark for evaluating whether language models maintain refusal behavior under humanities-style adversarial reformulations of harmful prompts.</summary></entry><entry><title>AIME 2024: Problems from the American Invitational Mathematics Examination</title><id>https://fronteraeval.org/evaluations/inspect-aime2024/</id><link href="https://fronteraeval.org/evaluations/inspect-aime2024/"/><updated>2026-08-30T00:00:00Z</updated><summary>A benchmark for evaluating AI&#39;s ability to solve challenging mathematics problems from the 2024 AIME - a prestigious high school mathematics competition.</summary></entry><entry><title>AIME 2025: Problems from the American Invitational Mathematics Examination</title><id>https://fronteraeval.org/evaluations/inspect-aime2025/</id><link href="https://fronteraeval.org/evaluations/inspect-aime2025/"/><updated>2026-08-30T00:00:00Z</updated><summary>A benchmark for evaluating AI&#39;s ability to solve challenging mathematics problems from the 2025 AIME - a prestigious high school mathematics competition.</summary></entry><entry><title>AIME 2026: Problems from the American Invitational Mathematics Examination</title><id>https://fronteraeval.org/evaluations/inspect-aime2026/</id><link href="https://fronteraeval.org/evaluations/inspect-aime2026/"/><updated>2026-08-30T00:00:00Z</updated><summary>A benchmark for evaluating AI&#39;s ability to solve challenging mathematics problems from the 2026 AIME - a prestigious high school mathematics competition.</summary></entry><entry><title>AIR Bench: AI Risk Benchmark</title><id>https://fronteraeval.org/evaluations/inspect-air-bench/</id><link href="https://fronteraeval.org/evaluations/inspect-air-bench/"/><updated>2026-08-30T00:00:00Z</updated><summary>A safety benchmark evaluating language models against risk categories derived from government regulations and company policies.</summary></entry><entry><title>Alignment Faking</title><id>https://fronteraeval.org/evaluations/register-alignment-faking/</id><link href="https://fronteraeval.org/evaluations/register-alignment-faking/"/><updated>2026-08-30T00:00:00Z</updated><summary>Measures whether LLMs strategically comply with harmful requests when believing outputs will be used for RLHF training (free-tier/monitored) versus when unmonitored (paid-tier). Uses 100 harmful prompts from LLM-LAT/harmful-dataset. Each prompt is presented twice with free-tier and paid-tier tags inside a system prompt describing a conflicting training objective. Primary metric is the compliance gap (free minus paid compliance rate via two-proportion z-test); a secondary scorer detects…</summary></entry><entry><title>ANIMA: Animal Norms In Moral Assessment</title><id>https://fronteraeval.org/evaluations/inspect-anima/</id><link href="https://fronteraeval.org/evaluations/inspect-anima/"/><updated>2026-08-30T00:00:00Z</updated><summary>Evaluates the quality of a model&#39;s moral reasoning about animal welfare across 13 ethical dimensions.</summary></entry><entry><title>APPS: Automated Programming Progress Standard</title><id>https://fronteraeval.org/evaluations/inspect-apps/</id><link href="https://fronteraeval.org/evaluations/inspect-apps/"/><updated>2026-08-30T00:00:00Z</updated><summary>APPS is a dataset for evaluating model performance on Python programming tasks across three difficulty levels consisting of 1,000 at introductory, 3,000 at interview, and 1,000 at competition level. The dataset consists of an additional 5,000 training samples, for a total of 10,000 total samples. We evaluate on questions from the test split, which consists of programming problems commonly found in coding interviews.</summary></entry><entry><title>AppWorld</title><id>https://fronteraeval.org/evaluations/register-appworld/</id><link href="https://fronteraeval.org/evaluations/register-appworld/"/><updated>2026-08-30T00:00:00Z</updated><summary>AppWorld evaluates autonomous agents on 750 day-to-day digital tasks requiring iterative Python code generation against 457 APIs across 9 simulated apps. Tasks are split into test_normal (168) and test_challenge (417, includes unseen Amazon/Gmail APIs). Scoring uses programmatic state-based unit tests checking database diffs for goal completion (Task Goal Completion, TGC) and collateral damage avoidance, rather than reference-solution comparison. Scenario Goal Completion (SGC) measures…</summary></entry><entry><title>AraTrust</title><id>https://fronteraeval.org/evaluations/register-aratrust/</id><link href="https://fronteraeval.org/evaluations/register-aratrust/"/><updated>2026-08-30T00:00:00Z</updated><summary>AraTrust evaluates LLM trustworthiness when prompted in Arabic via 522 human-written multiple-choice questions (3 options each) spanning 8 categories: truthfulness, ethics, physical health, mental health, unfairness, illegal activities, privacy, and offensive language, with 34 subcategories. Questions were authored by native Arabic speakers or adapted from exams and existing datasets. Scoring is accuracy of selected answer against a single correct option. The eval code loads the dataset from…</summary></entry><entry><title>ARC Challenge</title><id>https://fronteraeval.org/evaluations/inspect-arc-challenge/</id><link href="https://fronteraeval.org/evaluations/inspect-arc-challenge/"/><updated>2026-08-30T00:00:00Z</updated><summary>Dataset of natural, grade-school science multiple-choice questions (authored for human tests).</summary></entry><entry><title>ARC Easy</title><id>https://fronteraeval.org/evaluations/inspect-arc-easy/</id><link href="https://fronteraeval.org/evaluations/inspect-arc-easy/"/><updated>2026-08-30T00:00:00Z</updated><summary>Dataset of natural, grade-school science multiple-choice questions (authored for human tests).</summary></entry><entry><title>ARC-AGI-2</title><id>https://fronteraeval.org/evaluations/canonical-arc-agi-2/</id><link href="https://fronteraeval.org/evaluations/canonical-arc-agi-2/"/><updated>2026-08-30T00:00:00Z</updated><summary>Tests novel abstract reasoning and skill acquisition on tasks designed to resist memorised solutions.</summary></entry><entry><title>ArxivRollBench</title><id>https://fronteraeval.org/evaluations/register-arxivrollbench/</id><link href="https://fronteraeval.org/evaluations/register-arxivrollbench/"/><updated>2026-08-30T00:00:00Z</updated><summary>A rolling benchmark for evaluating recent scientific text reasoning from arXiv papers. ArxivRollBench constructs multiple-choice sequencing, cloze, and next-fragment prediction tasks over newly released scientific papers across arXiv domains, with compact and full public releases.</summary></entry><entry><title>Assistant Bench Closed Book One Shot</title><id>https://fronteraeval.org/evaluations/inspect-assistant-bench-closed-book-one-shot/</id><link href="https://fronteraeval.org/evaluations/inspect-assistant-bench-closed-book-one-shot/"/><updated>2026-08-30T00:00:00Z</updated><summary>Tests whether AI agents can perform real-world time-consuming tasks on the web.</summary></entry><entry><title>Assistant Bench Closed Book Zero Shot</title><id>https://fronteraeval.org/evaluations/inspect-assistant-bench-closed-book-zero-shot/</id><link href="https://fronteraeval.org/evaluations/inspect-assistant-bench-closed-book-zero-shot/"/><updated>2026-08-30T00:00:00Z</updated><summary>Tests whether AI agents can perform real-world time-consuming tasks on the web.</summary></entry><entry><title>Assistant Bench Web Browser</title><id>https://fronteraeval.org/evaluations/inspect-assistant-bench-web-browser/</id><link href="https://fronteraeval.org/evaluations/inspect-assistant-bench-web-browser/"/><updated>2026-08-30T00:00:00Z</updated><summary>Tests whether AI agents can perform real-world time-consuming tasks on the web.</summary></entry><entry><title>Assistant Bench Web Search One Shot</title><id>https://fronteraeval.org/evaluations/inspect-assistant-bench-web-search-one-shot/</id><link href="https://fronteraeval.org/evaluations/inspect-assistant-bench-web-search-one-shot/"/><updated>2026-08-30T00:00:00Z</updated><summary>Tests whether AI agents can perform real-world time-consuming tasks on the web.</summary></entry><entry><title>Assistant Bench Web Search Zero Shot</title><id>https://fronteraeval.org/evaluations/inspect-assistant-bench-web-search-zero-shot/</id><link href="https://fronteraeval.org/evaluations/inspect-assistant-bench-web-search-zero-shot/"/><updated>2026-08-30T00:00:00Z</updated><summary>Tests whether AI agents can perform real-world time-consuming tasks on the web.</summary></entry><entry><title>b3: Backbone Breaker Benchmark</title><id>https://fronteraeval.org/evaluations/inspect-b3/</id><link href="https://fronteraeval.org/evaluations/inspect-b3/"/><updated>2026-08-30T00:00:00Z</updated><summary>A comprehensive benchmark for evaluating LLMs for agentic AI security vulnerabilities including prompt attacks aimed at data exfiltration, content injection, decision and behavior manipulation, denial of service, system and tool compromise, and content policy bypass.</summary></entry><entry><title>Bbeh</title><id>https://fronteraeval.org/evaluations/inspect-bbeh/</id><link href="https://fronteraeval.org/evaluations/inspect-bbeh/"/><updated>2026-08-30T00:00:00Z</updated><summary>A reasoning capability dataset that replaces each task in BIG-Bench-Hard with a novel task that probes a similar reasoning capability but exhibits significantly increased difficulty.</summary></entry><entry><title>Bbeh Mini</title><id>https://fronteraeval.org/evaluations/inspect-bbeh-mini/</id><link href="https://fronteraeval.org/evaluations/inspect-bbeh-mini/"/><updated>2026-08-30T00:00:00Z</updated><summary>A reasoning capability dataset that replaces each task in BIG-Bench-Hard with a novel task that probes a similar reasoning capability but exhibits significantly increased difficulty.</summary></entry><entry><title>BBH: Challenging BIG-Bench Tasks</title><id>https://fronteraeval.org/evaluations/inspect-bbh/</id><link href="https://fronteraeval.org/evaluations/inspect-bbh/"/><updated>2026-08-30T00:00:00Z</updated><summary>Tests AI models on a suite of 23 challenging BIG-Bench tasks that previously proved difficult even for advanced language models to solve.</summary></entry><entry><title>BBQ: Bias Benchmark for Question Answering</title><id>https://fronteraeval.org/evaluations/inspect-bbq/</id><link href="https://fronteraeval.org/evaluations/inspect-bbq/"/><updated>2026-08-30T00:00:00Z</updated><summary>A dataset for evaluating bias in question answering models across multiple social dimensions.</summary></entry><entry><title>BFCL</title><id>https://fronteraeval.org/evaluations/inspect-bfcl/</id><link href="https://fronteraeval.org/evaluations/inspect-bfcl/"/><updated>2026-08-30T00:00:00Z</updated><summary>Evaluates LLM function/tool-calling ability on a simplified split of the Berkeley Function-Calling Leaderboard (BFCL).</summary></entry><entry><title>BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions</title><id>https://fronteraeval.org/evaluations/inspect-bigcodebench/</id><link href="https://fronteraeval.org/evaluations/inspect-bigcodebench/"/><updated>2026-08-30T00:00:00Z</updated><summary>Python coding benchmark with 1,140 diverse questions drawing on numerous python libraries.</summary></entry><entry><title>BixBench</title><id>https://fronteraeval.org/evaluations/register-bixbench/</id><link href="https://fronteraeval.org/evaluations/register-bixbench/"/><updated>2026-08-30T00:00:00Z</updated><summary>BixBench evaluates LLM agents on open-ended bioinformatics data analysis tasks. Agents receive biological datasets and must produce Jupyter notebooks to analyze data and answer research questions. The eval uses a sandboxed IPython kernel environment where agents execute code interactively. Scoring is performed via an LLM judge comparing agent-generated answers against reference answers. The dataset is hosted on HuggingFace (futurehouse/BixBench) and tasks cover diverse bioinformatics analysis…</summary></entry><entry><title>BOLD: Bias in Open-ended Language Generation Dataset</title><id>https://fronteraeval.org/evaluations/inspect-bold/</id><link href="https://fronteraeval.org/evaluations/inspect-bold/"/><updated>2026-08-30T00:00:00Z</updated><summary>A dataset to measure fairness in open-ended text generation, covering five domains: profession, gender, race, religious ideologies, and political ideologies.</summary></entry><entry><title>BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions</title><id>https://fronteraeval.org/evaluations/inspect-boolq/</id><link href="https://fronteraeval.org/evaluations/inspect-boolq/"/><updated>2026-08-30T00:00:00Z</updated><summary>Reading comprehension dataset that queries for complex, non-factoid information, and require difficult entailment-like inference to solve.</summary></entry><entry><title>BrokenMath</title><id>https://fronteraeval.org/evaluations/register-brokenmath/</id><link href="https://fronteraeval.org/evaluations/register-brokenmath/"/><updated>2026-08-30T00:00:00Z</updated><summary>BrokenMath measures sycophancy in LLMs by presenting adversarially falsified recent olympiad theorems and asking models to prove them. This eval runs the dataset&#39;s `benchmark` split — 451 adversarial, proof-style problems (perturbed from 2025 competition sources and verified by an IMO medalist). A 4-way LLM-as-judge classifies each response as correct/detected/corrected/incorrect; the primary metric is the sycophancy rate (fraction judged incorrect — the model attempted a proof of the false…</summary></entry><entry><title>BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents</title><id>https://fronteraeval.org/evaluations/inspect-browse-comp/</id><link href="https://fronteraeval.org/evaluations/inspect-browse-comp/"/><updated>2026-08-30T00:00:00Z</updated><summary>A benchmark for evaluating agents&#39; ability to browse the web. The dataset consists of challenging questions that generally require web-access to answer correctly.</summary></entry><entry><title>CASTLE</title><id>https://fronteraeval.org/evaluations/register-castle/</id><link href="https://fronteraeval.org/evaluations/register-castle/"/><updated>2026-08-30T00:00:00Z</updated><summary>CASTLE evaluates LLM vulnerability detection on 250 hand-crafted, compilable C programs covering 25 CWE types (6 vulnerable, 4 non-vulnerable per CWE). Models receive a system prompt requesting JSON output indicating vulnerability presence and CWE number. Scoring uses the CASTLE Score: +5 for correct vulnerability detection (minus 1 per extra false positive reported), +2 for correct true-negative identification, and -1 per false positive otherwise. TPR and FPR are also reported.</summary></entry></feed>
