evaluation-task · Source-linked discovery
SQuAD: A Reading Comprehension Benchmark requiring reasoning over Wikipedia articles
Set of 100,000+ questions posed by crowdworkers on a set of Wikipedia articles, where the answer to each question is a segment of text from the corresponding reading passage.
OriginPranav Rajpurkar, Jian Zhang, Konstantin Lopyrev et al.TopicsGeneral capabilityStatusimported
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.