Asking whether AI understands humor hides several questions inside one verb. Can a system recognize a joke? Explain a pun? Predict a preferred punchline? Generate a line for a specified audience? Infer a speaker’s intention? Experience amusement? Success on one task does not automatically answer the others.
Operational understanding
In engineering, “understands” sometimes means performs a defined task reliably. A classifier identifies sarcasm; a model explains the alternate meanings in a pun; a generator follows constraints. These are meaningful capabilities and can be tested against data and human judgments.
The problem begins when task performance is silently expanded into human-like comprehension. Humor depends on communicative intention, social relationship, culture, timing, and embodied situations. A benchmark samples some of that field, never all of it.
Form, meaning, and grounding
Language models learn powerful regularities in linguistic form. Bender and Koller argue for keeping form and grounded meaning distinct when describing language technology (Climbing Towards NLU). Their position is part of an active debate, but the methodological warning is clear: fluent continuation is not by itself evidence of human-analogous understanding.
A model may accurately explain that a joke uses two meanings because similar explanations occur in its learned patterns and the context supports them. That performance is useful. It does not reveal whether the system relates the language to a lived world as a person does.
Explanation is evidence with limits
Explanations expose more than yes/no labels. A strong explanation can identify setup, assumption, turn, required knowledge, and possible failure. Yet systems can produce plausible explanations for broken material. The 2023 exploratory study by Jentzsch and Kersting reported confident fictional explanations among its tested outputs (ACL Anthology).
An explanation should therefore be checked against the actual text. Fluency is not validation.
Humor includes disagreement
Even humans disagree about whether a line is sarcastic, what a joke targets, or why it works. That does not make evaluation impossible; it means evaluation should preserve distributions and context rather than invent one true label.
A 2025 survey of computational humor emphasizes subjective and ethical complexity alongside technical progress (Loakman and colleagues). A model may outperform some annotators on a narrow task while remaining unreliable across cultures, media, or unstated social situations.
What evidence would justify which claim?
- Recognition: performance across transparent, varied datasets.
- Explanation: accurate mechanisms and context, including uncertainty.
- Generation: human evaluation with originality and safety review.
- Adaptation: appropriate changes across audiences without stereotyping.
- Human-like understanding: a much broader theory of meaning, grounding, intention, and experience.
The final category cannot be inferred merely by accumulating the first four.
A productive middle position
It is unnecessary to deny impressive behavior in order to reject exaggerated conclusions. Systems can help writers, produce effective lines, and reveal patterns in comic language. They can also hallucinate context, flatten voice, and mistake a common template for a fresh observation.
ARTFunny studies those capabilities at the level the evidence supports. Read Does AI Understand Context? or inspect a machine specimen in the experiment laboratory.
