Skip to content

Latest commit

 

History

History
49 lines (29 loc) · 1.98 KB

File metadata and controls

49 lines (29 loc) · 1.98 KB

Task-name

Paper

Title: Global PIQA

Abstract: To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. We present Global PIQA, a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by 320 researchers from 65 countries around the world. The 116 language varieties in Global PIQA cover five continents, 14 language families, and 23 writing systems. In the non-parallel split of Global PIQA, over 50% of examples reference local foods, customs, traditions, or other culturally-specific elements. Beyond its uses for LLM evaluation, we hope that Global PIQA provides a glimpse into the wide diversity of cultures in which human language is embedded.

Short description of paper / benchmark goes here:

Homepage: homepage to the benchmark's website goes here, if applicable

Citation

BibTeX-formatted citation goes here

Groups, Tags, and Tasks

Groups

  • group_name: global_piqa_completions Generation task using chat format
  • group_name: global_piqa_prompted Cloze-style completion format

Tags

  • tag_name: Short description

Tasks

  • task_name: 1-sentence description of what this particular task does
  • task_name2: ...

Checklist

For adding novel benchmarks/datasets to the library:

  • Is the task an existing benchmark in the literature?
    • Have you referenced the original paper that introduced the task?
    • If yes, does the original paper provide a reference implementation? If so, have you checked against the reference implementation and documented how to run such a test?

If other tasks on this dataset are already supported:

  • Is the "Main" variant of this task clearly denoted?
  • Have you provided a short sentence in a README on what each new variant adds / evaluates?
  • Have you noted which, if any, published evaluation setups are matched by this variant?

Changelog