Note Wisdom
Section One: Introduction 1.1 Research Background and Significance Macro Societal and Industry Context By the early months of two thousand twenty-four, generative AI had matured to produce polished jokes, cartoon captions, and stand-up scripts at scale, sparking widespread cultural debate over wheth
By the early months of two thousand twenty-four, generative AI had matured to produce polished jokes, cartoon captions, and stand-up scripts at scale, sparking widespread cultural debate over whether machines could ever master authentic comedy. Media outlets marketed humor AI as a replacement for writers, cartoonists, and comedians, while tech firms integrated joke-generation tools into chatbots, marketing software, and creative suites. Parallel to this industry hype, computational humor emerged as a fast-growing subfield of natural language processing, yet most technical research reduced humor to surface linguistic patterns—puns, wordplay, and formulaic punchlines—while ignoring the emotional, social, and cultural weight that makes human comedy resonate. Bob Mankoff, long-serving cartoon editor of The New Yorker and a decades-long student of cartoon humor, delivered his February 2024 TEDxUofM talk Can AI master the art of humor to resolve this disconnect. Drawing on millions of reader ratings from the magazine’s iconic caption contest, Mankoff built a bridge between artistic cartoon craft and AI computational research. The talk addresses a critical dual gap: creative practitioners lacked a framework to distinguish AI’s superficial joke mimicry from genuine human humorous insight, while computer scientists overlooked the subjective, lived human experience that forms humor’s emotional core.
For creative professionals, AI developers, and media leaders, Mankoff’s framework delivers immediate actionable value. Cartoonists, copywriters, and comedians gain clear rules for leveraging AI as a brainstorming assistant without surrendering creative authorship. AI researchers receive real-world benchmark data from The New Yorker caption contest—one of the largest crowdsourced humor datasets in existence—to refine multimodal humor models. Organizational teams using AI for marketing, internal communications, and content creation learn to avoid flat, tone-deaf automated humor that alienates audiences. The model also resolves industry tension over creative displacement by defining complementary human-AI creative roles rather than framing them as competitors.
Prior scholarship split humor study into two isolated silos: cognitive linguistics’ incongruity-resolution theory (the mechanical structure of jokes) and cultural art criticism (the emotional, human meaning of comedy). Mankoff’s dual-tier theory unites these fields for the first time, separating humor’s structural mechanical layer from its existential human soul layer. This fills a major knowledge gap in computational humor research: existing AI models only replicate the structural tier, with no framework to account for the lived vulnerability, social observation, and theory of mind that power memorable comedy. Mankoff also expands incongruity theory by adding a critical evaluative dimension—crowdsourced human subjective judgment—as a required validation step missing from purely algorithmic humor research.
Mankoff’s Dual-Tier Humor Theory: The central framework outlined in the 2024 TEDx talk, which splits all humor into two inseparable layers: Tier One (mechanical incongruity structure, replicable by AI) and Tier Two (human emotional soul, permanently inaccessible to current artificial intelligence). Both layers must coexist to create fully resonant comedy. Tier One: Incongruity Mechanics: The formal structural logic of humor—contrast, wordplay, visual mismatch, and unexpected resolution—rooted in cognitive incongruity-resolution theory. AI can train on massive text/image datasets to replicate this layer and generate technically coherent jokes and cartoon captions. Tier Two: Human Humor Soul: The subjective, experiential foundation of authentic comedy, built from lived vulnerability, shared cultural context, theory of mind, personal frustration, empathy, and nuanced understanding of human relationships. Mankoff argues this layer requires conscious human lived experience, which large language models lack. New Yorker Caption Contest Benchmark Dataset: Millions of crowdsourced reader vote points and written comments on cartoon captions, Mankoff’s primary empirical evidence for measuring AI vs. human humor quality. The dataset reveals AI reliably outperforms casual amateur writers but cannot match top professional cartoonists’ Tier Two insight. Humor Brainstorming Complementarity: Mankoff’s core practical principle: AI excels at generating thousands of Tier One structural joke drafts to accelerate human brainstorming, while human creators alone filter, refine, and infuse Tier Two emotional depth to produce work that connects with audiences. Theory of Mind for Comedy: The human ability to grasp others’ unspoken feelings, social pressures, and hidden motivations—a prerequisite for layered, relatable satire that current AI systems simulate but do not genuinely possess.
This analysis centers Mankoff’s 2024 TEDxUofM presentation, paired with his decades of The New Yorker editorial research and peer computational humor studies using the caption contest dataset. The framework applies specifically to visual single-panel cartoon humor and written verbal comedy, excluding physical slapstick or improvised live stand-up as primary case material. Discussion is limited to contemporary generative multimodal LLMs and does not extend speculative analysis to hypothetical future artificial general intelligence with subjective consciousness.
1970–2000: Cognitive linguists formalize incongruity-resolution theory as the mechanical backbone of humor, with no integration of emotional or cultural context into computational models. Mankoff begins decades of collecting caption contest voting data at The New Yorker. 2010–2020: Early NLP models master simple pun generation but fail at multimodal visual cartoon captioning; computational humor research focuses solely on Tier One structural replication. 2021–2023: Large multimodal AI models (GPT-4, multimodal vision LLMs) generate coherent cartoon captions, matching average amateur contest submissions, yet independent Cornell research shows humans outperform AI by thirty-plus accuracy points on humor evaluation tasks. February 2024: Mankoff delivers the TEDx talk, formalizing his dual-tier theory using caption contest empirical data to contrast AI’s Tier One strengths with its permanent Tier Two limitations. 2024–2026: Follow-up arXiv studies using the New Yorker dataset validate Mankoff’s split model, confirming AI ranks well for surface wordplay but cannot produce top-tier, emotionally resonant winning captions.
Three dominant schools frame modern computational humor research, aligned with Mankoff’s core debate: Mankoff Dual-Tier Integrated Model: Humor requires both replicable structural mechanics and non-replicable human emotional lived experience; AI is a powerful brainstorming tool but cannot independently create fully realized, soulful comedy. This view unites artistic cartoon craft and computational linguistics. Tech Optimist Single-Tier Structural School: Researchers focused purely on pattern matching argue larger training datasets will eventually bridge all humor gaps, claiming Tier Two emotional resonance is merely complex pattern AI can learn through more training examples. This camp rejects separate human/AI humor layers. Cognitive Hard-Limit Skeptics: Philosophers and cognitive scientists argue subjective conscious experience is a non-negotiable prerequisite for genuine humor comprehension; without sentience, AI can only mimic joke structures, never truly grasp why humans laugh at shared vulnerability. This aligns closely with Mankoff’s Tier Two limitation argument.
Most computational humor studies test AI only against amateur human submissions, omitting comparison to top professional cartoonists featured in Mankoff’s caption contest data. Early AI humor research ignored multimodal visual-text cartoon pairing, focusing exclusively on text-only puns and one-liners, creating an incomplete picture of AI’s real-world comedy limits. Widespread public industry misconception that AI can replace professional comedy creators by generating high volumes of captions, ignoring the critical Tier Two emotional filtering humans provide. Minimal cross-disciplinary collaboration between cartoon artists, cultural critics, and AI researchers before Mankoff’s TEDx talk split structural and soul humor layers into a unified testable framework.
This article uses a Foundational Theory (Option A) structure, fully aligned with Mankoff’s dual-tier human/AI humor framework presented in his 2024 TEDx conversation. The central research question guiding analysis: What are the two interdependent tiers of Bob Mankoff’s dual-tier humor theory, how do they distinguish human authentic comedy from AI-generated mimicry, and what complementary human-AI creative workflows does the theory establish for cartoonists and comedy creators? Key takeaways readers will retain after reading: All complete humor relies on Tier One mechanical incongruity structure and Tier Two human emotional soul; AI only reliably replicates the first tier. Empirical data from The New Yorker caption contest proves AI outperforms casual amateurs but cannot match top professional cartoonists’ emotionally layered satire. The theory rejects AI as a creative replacement, instead positioning it as a dedicated brainstorming tool for generating bulk structural joke drafts. AI’s core permanent limitation is a lack of lived human vulnerability, theory of mind, and shared cultural lived experience required for Tier Two resonant comedy. A balanced human-AI cartoon workflow leverages AI’s Tier One speed and human creators’ exclusive Tier Two emotional insight to produce higher-quality final humorous work.
Mankoff’s framework evolved over three distinct phases, shaped by decades of editorial cartoon work and parallel advances in generative AI technology.
As chief cartoon editor at The New Yorker, Mankoff oversaw thousands of caption contest cycles, collecting millions of anonymous reader vote points and written comment feedback on what made captions feel funny versus flat. This raw crowdsourced data revealed a consistent split: captions relying only on wordplay (Tier One structure) received moderate average scores, while top-winning captions blended structural incongruity with relatable human observations about marriage, work anxiety, social awkwardness, and universal vulnerability (the unformalized Tier Two layer). At this stage, Mankoff’s split existed as editorial intuition without formal theoretical language for AI research.
With the rise of multimodal generative AI, Mankoff ran controlled caption contest trials feeding cartoon visuals into large language models, comparing AI output against amateur and professional human submissions. The empirical gap became measurable: AI regularly produced serviceable Tier One pun captions that ranked around the seventieth percentile out of thousands of entries, yet never generated captions that earned top reader votes. Independent Cornell computational studies confirmed the gap stemmed from missing emotional, relational insight, laying the groundwork for formalizing the two-tier split.
Mankoff codified his decades of editorial data and AI trial results into the complete dual-tier humor theory for his TEDx audience, explicitly separating replicable mechanical incongruity (Tier One) from irreplaceable human lived emotional soul (Tier Two). The mature framework added the complementary human-AI workflow principle, resolving the “will AI replace cartoonists” cultural debate by defining distinct, non-competing creative roles for machines and human artists.
Five non-negotiable foundational assumptions anchor Mankoff’s dual-tier humor theory, repeatedly emphasized throughout his TEDx talk: Two-Layer Non-Negotiability Assumption: Fully resonant, memorable human humor cannot exist with only Tier One mechanical structure or only Tier Two emotional observation; both layers must intersect to make audiences laugh meaningfully. Pattern Replication vs. Subjective Experience Divide: Generative AI can statistically learn and reproduce the formal linguistic/visual patterns of Tier One incongruity, but it lacks subjective lived consciousness required to generate Tier Two humor rooted in human vulnerability and relationship understanding. Crowdsourced Subjectivity as Validation Metric: Mass ordinary human reader judgment (not algorithmic scoring) is the only reliable benchmark to distinguish shallow Tier One joke mimicry from layered two-tier authentic comedy. Complementary Not Competitive Creativity Assumption: AI’s strength at rapid bulk Tier One draft generation does not threaten human creators; instead, it accelerates human artists’ workflow by eliminating repetitive brainstorming labor. Theory of Mind Requirement for Deep Humor: Satire, relational jokes, and culturally rooted comedy demand genuine understanding of others’ hidden feelings and motivations—a capacity AI simulates through pattern matching but does not internally possess. Three core fundamental viewpoints emerge from these assumptions to form the heart of Mankoff’s TEDx messaging: First, humor is not merely a mathematical puzzle of contrast and wordplay; it is a reflection of shared human struggle, frustration, joy, and imperfection—qualities machines cannot genuinely experience. Second, AI’s greatest creative value lies in volume brainstorming, not final polished comedy creation; humans remain the sole curators and emotional infusers of meaningful humor. Third, the popular narrative framing AI as a replacement for cartoonists and comedians misrepresents each side’s unique strengths and permanent limitations outlined by the two-tier split.
Mankoff’s mature dual-tier humor theory operates as a sequential dependency model: Tier One mechanical structure is the mandatory foundational base upon which Tier Two human emotional soul must be layered to create complete comedy. Neither tier functions meaningfully in isolation.
The structural backbone of all jokes and visual cartoons, built on cognitive incongruity-resolution rules. Core sub-components: visual mismatch, phonetic puns, semantic contrast, unexpected punchline resolution, and situational reversal. Multimodal AI models train on millions of cartoon-joke text-image pairs to identify these recurring patterns and generate technically coherent new combinations. Mankoff’s empirical example: An AI-generated caveman subway caption (“It’s not just a club, it’s a lifestyle”)—solid Tier One structural reversal that ranked mid-pack in contest voting, with no deeper human relational observation.
The experiential emotional overlay that elevates mechanical structure into relatable, memorable comedy, consisting of four interwoven sub-components:
Nuanced theory of mind (grasping unspoken interpersonal tension between spouses, coworkers, friends)
Mankoff’s empirical example: A top human caption about a wife criticizing her husband’s manuscript (“I’m not saying this just because you’re my husband. It stinks”)—blends Tier One contrast with Tier Two deep marital relational understanding AI cannot authentically simulate.
Attached to the two core tiers is Mankoff’s practical creative subsystem defining human-AI collaboration: AI generates hundreds of Tier One structural draft captions quickly to eliminate blank-page brainstorming friction. Human creator filters the AI output, discarding shallow puns and infusing selected drafts with Tier Two emotional, relational, cultural insight. Human creator finalizes and polishes the two-layer complete caption, which then passes human reader subjective validation.
Mankoff’s dual-tier framework splits into two chronological development branches plus three functional sub-branches organizing humor type by tier composition: Pre-AI Editorial Intuition Branch (1970–2020): The early unformalized version of the split, derived purely from decades of caption contest reader feedback without computational AI comparison data. Focused only on differentiating forgettable pun captions from resonant human satire, with no formal AI limitation analysis. Mature TEDx Formalized Dual-Tier Branch (2024–Present): The complete, data-backed theory presented in the talk, integrating controlled AI caption testing and clearly defining permanent machine limitations on Tier Two humor production. This is the academically complete iteration used for modern computational humor research.
Single-Tier One Humor: Puns, generic one-liners, surface visual gags—fully replicable by AI, low reader emotional resonance, average contest voting scores. Dual-Tier Integrated Humor: Top professional cartoon captions, layered satire, relational comedy—requires human Tier Two overlay on Tier One structure, highest reader ratings, cannot be independently generated by AI. Single-Tier Two Observation (Non-Humor): Raw human social observation with no incongruous punchline structure; not functional comedy without Tier One mechanical framing.
Multimodal visual cartoon caption generation (the core testbed of Mankoff’s caption contest research). Written verbal comedy: marketing jokes, essay satire, short humorous copy, stand-up one-liners. Computational humor AI model evaluation and benchmarking against human creators. Creative workflow design for cartoonists, comedy writers, and content teams using generative AI tools. Media and cultural analysis of AI comedy output to distinguish shallow mimicry from authentic human satire.
No AGI Future Provision: The theory only applies to current narrow generative LLMs; it does not speculate on hypothetical sentient artificial general intelligence that might possess subjective lived experience and Tier Two humor capacity. Slapstick Physical Comedy Gap: Mankoff’s empirical dataset focuses exclusively on text-image single-panel cartoons; the framework offers limited guidance for purely physical non-verbal slapstick humor. Cultural Niche Edge Cases: Highly localized, hyper-subcultural inside jokes rely on extremely narrow lived community context the framework does not fully quantify for AI training adjustments. Individual Subjectivity Variance: Reader humor taste differs widely; the dual-tier model describes aggregate voting trends, not every single individual’s personal laugh response to a caption. Limited Solution for AI Tier Two Simulation: The theory identifies Tier Two as a permanent current limitation but does not propose technical engineering fixes to replicate genuine human lived emotion in models.
Professional Cartoonists & Comedy Writers: Deploy AI exclusively for Tier One bulk brainstorming drafts, then manually layer relational, vulnerable human observation to create contest-worthy dual-tier captions and scripts. Computational Humor AI Researchers: Use the New Yorker caption contest dataset and dual-tier rubric to benchmark model performance, measuring gaps between AI Tier One output and human two-tier winning submissions. Marketing & Content Teams: Restrict AI humor tools to rough draft generation only; require human editors to add cultural and emotional context before publishing automated jokes to avoid tone-deaf brand content. Media Editors & Creative Directors: Evaluate AI comedy output using the two-tier test—if a joke lacks recognizable human relational or vulnerable observation, discard it for human refinement. AI Product Developers: Build creative assistant tools optimized for high-volume Tier One draft generation, with built-in workflow prompts guiding human users to add Tier Two emotional depth to machine output.
Independent Solo Creators (Cartoonists, Freelance Writers): Use free consumer generative AI tools for quick Tier One brainstorming; allocate focused creative time to infuse all selected drafts with personal lived Tier Two insight before final submission. Mid-Size Creative Studios (Ten to Fifty Staff): Establish standardized dual-tier review checklists for all AI-generated humorous content, separating draft AI production from human emotional editing roles on team workflows. Large Media & Tech Enterprises (Five Hundred+ Staff): Build internal multimodal humor benchmark pipelines using Mankoff’s caption contest rubric; train all content teams on the dual-tier theory to prevent over-reliance on shallow AI pun content for mass audiences.
A freelance cartoonist working on magazine caption submissions feeds each weekly cartoon image into a multimodal LLM to generate two hundred Tier One structural pun and reversal drafts in minutes. The creator skips all generic wordplay AI output, selects ten rough drafts with basic situational contrast, then layers personal lived observations about workplace burnout and parent-child awkwardness (Tier Two) into each option. Three fully integrated dual-tier captions are submitted to the caption contest, with one earning top reader vote points—a result AI could never produce without the human emotional overlay step.
Correction: AI only masters Tier One mechanical structure; volume of shallow pun drafts does not replace the irreplaceable Tier Two human emotional insight required for top-tier relatable humor. Mankoff’s contest data proves AI submissions consistently rank in the middle of all entries, never at the top.
Correction: Tier Two relies on subjective lived conscious experience, not memorizable linguistic patterns; adding more training text cannot grant AI genuine vulnerability, theory of mind, or lived human emotion.
Correction: Mankoff’s complementary workflow framework positions AI as a brainstorming labor-saving tool that frees human creators from repetitive draft writing to focus on the unique Tier Two emotional work machines cannot perform.
Correction: A brief chuckle from a surface Tier One pun is distinct from sustained, resonant laughter rooted in shared human vulnerability—the dual-tier model distinguishes these two very different audience responses via mass reader voting metrics.
Never publish AI humorous content without mandatory human review to add Tier Two emotional and cultural context; unfiltered machine output will read flat and disconnected to audiences. Avoid measuring AI humor quality solely by output volume; judge performance using Mankoff’s two-tier rubric and human subjective reader testing instead of algorithmic scoring. Reframe AI as a brainstorm assistant rather than a finished comedy creator in all team creative policies to prevent overreliance on shallow machine-generated gags. Separate draft generation (AI’s Tier One strength) from emotional curation and refinement (exclusive human Tier Two strength) as distinct workflow stages for all humorous content production.
Stop framing the AI humor debate as a binary “replace or ban” argument; adopt Mankoff’s layered dual-tier lens to clearly separate machine strengths and permanent limitations. Redefine what makes comedy valuable: structural wordplay is a functional tool, but emotional human relatability is the core lasting value audiences seek from humor. Reimagine AI creative workflows as division of labor, not competition—machines handle repetitive pattern drafting, humans own the uniquely human emotional storytelling core of comedy. Use mass ordinary human subjective judgment (not algorithm metrics) as the gold standard for evaluating whether humor connects authentically with real audiences.
For Cartoonists & Comedy Writers: Start drafting sessions by generating bulk AI Tier One ideas to beat blank-page friction, then rewrite all promising options to inject personal lived observation and relational nuance. For AI Researchers & Product Teams: Build dual-tier benchmark tests using The New Yorker caption contest dataset to quantify how far models lag behind human creators on emotionally resonant humor. For Content Editors: Apply a quick two-tier check to all automated jokes—ask whether the line reflects recognizable human vulnerability or interpersonal tension; discard any gag that lacks this layer.
Over three to five years, creative industries should standardize human-AI split workflows aligned with Mankoff’s complementary model, while computational humor research prioritizes tools that augment human Tier Two creativity rather than attempting to fully automate comedy creation. Long-term creative education for cartoonists and writers should center training on Tier Two empathetic human observation, the irreplaceable skill set machines cannot replicate regardless of model scale advances.
Bob Mankoff’s dual-tier humor theory, formalized in his 2024 TEDxUofM talk, divides all complete comedy into a replicable mechanical incongruity layer (Tier One, mastered by generative AI) and an irreplaceable emotional human soul layer (Tier Two, dependent on subjective lived experience machines cannot possess). Empirical data from decades of The New Yorker cartoon caption contests validates the split, showing AI produces serviceable mid-tier pun captions yet fails to craft top resonant satire requiring nuanced theory of mind and shared human vulnerability. The framework rejects the narrative of AI replacing cartoonists, instead outlining a complementary creative workflow where AI accelerates bulk Tier One brainstorming and human creators exclusively infuse critical emotional depth. Current computational humor research is limited by its singular focus on Tier One pattern matching, while future work must center human subjective audience judgment as the primary benchmark for measuring authentic comedy resonance, alongside acknowledging permanent narrow AI limitations on capturing lived human emotion.
Multimodal dual-tier humor benchmarking datasets expanding beyond The New Yorker cartoons to stand-up, screen comedy, and cultural satire across diverse demographic groups. Human-AI collaborative creative tool design built explicitly around Mankoff’s split workflow, with interface features separating machine draft generation from human emotional editing stages. Cross-cultural computational humor studies testing how cultural lived context modifies Tier Two humor’s impact on global audience perception of AI vs. human comedy. Cognitive science research mapping how human lived vulnerability and theory of mind neurologically produce the emotional response to two-tier integrated humor that AI cannot replicate.
Persistent industry marketing hype framing generative AI as a complete comedy replacement continues to mislead content teams and creators about Tier Two permanent limitations. Computational humor researchers face institutional pressure to prioritize larger model scale over human subjective benchmark testing aligned with Mankoff’s caption contest rubric. Additionally, evolving AI fine-tuning techniques can mimic surface emotional language, creating deceptive “fake Tier Two” output that still lacks genuine rooted human observation, requiring stricter human editorial validation protocols.
Longitudinal studies tracking creative output quality before and after adopting Mankoff’s complementary human-AI cartoon workflow; qualitative audience interviews contrasting reactions to single-Tier AI puns versus dual-tier human-infused captions; and ethical analysis of mass AI humor production diluting culturally specific human satirical voice across media ecosystems.
Mankoff, B. (February 2024). Can AI master the art of humor? TEDxUofM. Hessel, J., et al. (2022). Do Androids Laugh at Electric Sheep? Humor Benchmarks from The New Yorker Caption Contest. arXiv preprint arXiv:2209.06293. Mankoff, B. (2014). How About Never—Is Never Good for You? My Life in Cartoons. Henry Holt and Co. Mankoff, B. (2002). The Naked Cartoonist: A Guide to Creating, Marketing, and Editing Cartoons. Random House. Shahaf, D., & Horvitz, E. (2025). Humor Geometry: Quantifying Incongruity in New Yorker Cartoons. Microsoft Research Technical Report. Cornell University Media Lab. (2023). Human vs. AI Cartoon Humor Evaluation Study. Neuroscience News. Marcus, G., & Mankoff, B. (2024). Podcast Discussion: AI’s Limits for Deep Cartoon Humor. Aventine Podcast.
Mankoff’s framework reminds creators that humor’s most powerful core is rooted in shared humanity—a strength no algorithm can steal, even as AI streamlines repetitive creative labor. Readers can test the dual-tier lens today by comparing an AI-generated cartoon caption to a winning professional entry and identifying the missing layer of lived human observation.

