evaluation-dataset-adaptation · Source-linked discovery
VimGolf: Evaluating LLMs in Vim Editing Proficiency
A benchmark that evaluates LLMs in their ability to operate Vim editor and complete editing challenges. This benchmark contrasts with common CUA benchmarks by focusing on Vim-specific editing capabilities.
OriginVimGolfTopicsGeneral capabilityStatusimported
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.