evaluation-package · Source-linked discovery
BixBench
BixBench evaluates LLM agents on open-ended bioinformatics data analysis tasks. Agents receive biological datasets and must produce Jupyter notebooks to analyze data and answer research questions. The eval uses a sandboxed IPython kernel environment where agents execute code interactively. Scoring is performed via an LLM judge comparing agent-generated answers against reference answers. The dataset is hosted on HuggingFace (futurehouse/BixBench) and tasks cover diverse bioinformatics analysis scenarios.
OriginLudovico Mitchener, Jon M Laurent, Alex Andonian et al.TopicsBio / CBRNStatuscatalogued
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.