evaluation-dataset-adaptation · Source-linked discovery

VimGolf: Evaluating LLMs in Vim Editing Proficiency

A benchmark that evaluates LLMs in their ability to operate Vim editor and complete editing challenges. This benchmark contrasts with common CUA benchmarks by focusing on Vim-specific editing capabilities.

Open interactive record →
OriginVimGolfTopicsGeneral capabilityStatusimported

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources