Multilingual Concreteness Norms for Verb-Object Expressions

Human concreteness ratings and figurative or literal judgements for verb-object expressions in English, German and Slovene, such as throw ball and carry implication, with 9,000 example sentences written by native speakers.

  1. Institute for Natural Language Processing, University of Stuttgart
  2. Center for Information and Language Processing, LMU Munich
  3. School of Computation, Information and Technology, TU Munich

Download

Concreteness norms

Human ratings from 1 (abstract) to 5 (concrete) for 5,814 verb-object expressions, the first human-rated Slovene word norms, and model predictions for a much larger set.

Format
ZIP
Languages
English, German, Slovene
Download concreteness norms

Figurative language judgements

Figurative or literal labels for 1,800 verb-object expressions, with 9,000 example sentences written by native speakers.

Format
ZIP
Languages
English, German, Slovene
Download figurative judgements

Each archive holds tab-separated UTF-8 files and a README describing every column.

What is in the data

Verb-object expressions in English, German and Slovene placed on a scale from abstract to concrete. Each expression is shown with the individual concreteness scores of its verb and noun, and the score of the expression as a whole often sits well away from the average of the two.
An expression is not the average of its words: carry implication is abstract although carry is concrete. Dashed outlines mark figurative expressions.

Concreteness norms for verb-object expressions and Slovene words

Files in the concreteness archive.
File Rows Contents
vo_expressions/english.tsv 1,998 Human ratings for verb-object expressions, 5,814 in total
vo_expressions/german.tsv 1,818
vo_expressions/slovene.tsv 1,998
vo_expressions/english_rating_scale.tsv 1,823 The same English expressions rated a second time on a 1 to 5 scale, for comparing the two methods
words/slovene.tsv 798 First of its kind for Slovene
Human-rated Slovene word concreteness norms for 600 nouns and 198 verbs
vo_generated/english.tsv 470,856 Model-predicted ratings for every other expression extracted from the corpora, 613,602 in total
vo_generated/german.tsv 56,406
vo_generated/slovene.tsv 86,340

Figurative language judgements

For 600 expressions per language, five native speakers each judged the expression figurative or literal and wrote a sentence using it. Together this makes a multilingual resource for figurative language and metaphor research.

Files in the figurative archive.
File Rows One row per
judgements.tsv 1,800 Expression, with the majority decision and its concreteness score
example_sentences.tsv 9,000 Sentence, with the decision made by the person who wrote it

Labels need four of five annotators to agree; a 3 to 2 split is unsure.

Two rows as they appear in the data, with one of their five sentences.
Expression Label Concreteness Example sentence
carry uncertainty figurative 2.20 She carried uncertainty when faced with a difficult problem.
buy oil literal 4.35 I went to the supermarket to buy some oil

Research context

Literally Concrete or Figuratively Abstract? Multilingual Concreteness Norms for Verb-Object Expressions

Transactions of the Association for Computational Linguistics, volume 14, pages 1205–1224, 2026.

Concreteness norms usually rate single words. We rated verb-object expressions instead, 5,814 of them across three languages, extended them to a much larger set with automatic methods, and had a subset judged figurative or literal.

The noun matters more than the verb

Box plots of concreteness scores for the six verb-object groups in English, German and Slovene. Groups with abstract nouns sit near the bottom of the scale and groups with concrete nouns near the top, with the same ordering in all three languages.
Abstract nouns keep an expression low whatever the verb; concrete nouns lift it. The pattern holds in all three languages.

Abstract expressions read as figurative

Stacked bars showing the share of figurative, unclear and literal majority labels for each of the six concreteness groups, per language. The figurative share is largest for the most abstract groups and nearly disappears for the most concrete ones.
The more concrete the expression, the less often it is judged figurative. Note the large share of unsure labels, shown as Unclear.

Best-Worst Scaling is more reliable than a rating scale

Side-by-side distributions of English concreteness scores from Best-Worst Scaling and from a 1 to 5 rating scale, per concreteness group. The rating scale values bunch toward the middle of the range while the Best-Worst values spread across it.
English scores from Best-Worst Scaling (left) and a 1 to 5 rating scale (right). Rating-scale scores cluster in the middle and agree less between annotators. They are in english_rating_scale.tsv.

How good the predicted ratings are

We compared a baseline, fine-tuned monolingual models and prompted LLMs (monolingual and multilingual, zero- and few-shot). The fine-tuned models did best and produced the released ratings. The full comparison is in the paper.

Language Model Spearman ρ RMSE Baseline ρ
EnglishRoBERTa.880.38.75
GermanGBERT.870.45.77
SloveneSloBERTa.820.54.49

Cite this work

Urban Knupleš, Diego Frassinelli, Alexander Fraser, and Sabine Schulte im Walde. 2026. Literally Concrete or Figuratively Abstract? Multilingual Concreteness Norms for Verb-Object Expressions. Transactions of the Association for Computational Linguistics, 14:1205–1224.

@article{knuples-etal-2026-concreteness,
    title     = {Literally Concrete or Figuratively Abstract? {M}ultilingual
                 Concreteness Norms for Verb-Object Expressions},
    author    = {Knuple{\v{s}}, Urban and Frassinelli, Diego and
                 Fraser, Alexander and {Schulte im Walde}, Sabine},
    journal   = {Transactions of the Association for Computational Linguistics},
    volume    = {14},
    pages     = {1205--1224},
    year      = {2026},
    doi       = {10.1162/tacl.a.703}
}