When Cow Urine Cures Constipation on YouTube: Limits of LLMs in Detecting Culture-specific Health Misinformation
An LLM-assisted analysis of 30 multilingual YouTube videos shows that cultural health misinformation defeats prompt-only detection.
Key Findings
Methodology
The authors queried the YouTube Data API v3 with cow urine, gomutra, #gomata, and related terms, retaining 30 English, Hindi, or Urdu videos. Audio was transcribed with Whisper large, yielding 7.04% WER on a 16% manually checked sample, and non-English text was translated with GPT-4. Human labels covered stance and speaker gender. GPT-4o extracted traditional metaphors and scientific terms; GPT-4o-mini, Gemini 2.5 Pro, and DeepSeek-V3.1 extracted intensifiers under formal/friendly and zero-/few-shot prompts.
Key Results
- Twenty-four videos (80%) promoted gomutra and six (20%) debunked it. Promoting videos contained about 4.4 traditional metaphors and 2.8 scientific terms per 100 words, versus about 1.0 and 1.5 in debunking videos, showing that both cultural and scientific registers structure the discourse.
- GPT-4o achieved 61% precision, 53% recall, and 52% F1 for scientific terms; traditional-term scores were 64%, 56%, and 59%. This is a usable exploratory baseline, not dependable cultural interpretation.
- Intensifier agreement was poor: Gemini versus GPT-4o-mini κ=−0.256, Gemini versus DeepSeek-V3.1 κ=−0.006, and GPT-4o-mini versus DeepSeek-V3.1 κ=0.214. Formal versus friendly prompts yielded κ=0.431, while zero-shot versus few-shot yielded κ=0.452.
Significance
The paper shows that misinformation may persuade through religious identity, tradition, institutional authority, testimony, and scientific-sounding language rather than through easily isolated false propositions. Debunkers often reuse culturally meaningful terms to reach audiences, so a detector that treats traditional language as a misinformation cue can flag effective counter-speech. This matters for platform governance, public-health communication, and equitable evaluation in India and other Global South settings.
Technical Contribution
The study introduces a preliminary multilingual gomutra discourse corpus and combines discourse analysis, LLM extraction, prompt-sensitivity testing, and Cohen’s kappa. It evaluates not only whether models find lexical markers, but also how persona, shot setting, and gendered rhetoric affect measurement. The central engineering contribution is methodological: LLMs are framed as auditable, human-validated instruments rather than autonomous truth classifiers.
Novelty
According to the authors, this is the first systematic study of Indian multilingual cow-urine health discourse together with an evaluation of LLM cultural sensitivity. Unlike work centered on COVID-19 diffusion or binary truth labels, it jointly examines traditional metaphors, pseudo-scientific vocabulary, gendered persuasion, and prompt design.
Limitations
- The corpus contains only 30 videos, with an 80/20 promoting–debunking imbalance. A single author performed the main annotations, limiting representativeness, inter-annotator reliability, and statistical power.
- Only English, Hindi, and Urdu were retained; GPT translation, one native-speaker annotator, API relevance ranking, and changing model versions may introduce additional bias.
- Lexical and rhetorical counts do not establish medical truth, causal influence, or audience uptake.
Future Work
Future research should expand languages, regions, platforms, and time periods; use multi-expert cultural and medical annotation; and combine retrieval from evidence graphs with local models and multimodal signals. Preregistered prompts, repeated model runs, and human–LLM studies should separate cultural adaptation from factual verification and fairness.
AI Executive Summary
Social media is now a major health-information channel in the Global South. Indian YouTube videos about gomutra—cow urine—illustrate why conventional misinformation detection struggles: claims about constipation, obesity, detoxification, or cancer are wrapped in sacred tradition, institutional authority, testimony, and scientific-sounding vocabulary. Their persuasive force is cultural as much as factual.
Khan and colleagues collected 30 English, Hindi, and Urdu videos. Whisper large transcribed the audio with a 7.04% sampled word-error rate; GPT-4 translated non-English transcripts. GPT-4o identified traditional metaphors and scientific terms, while GPT-4o-mini, Gemini 2.5 Pro, and DeepSeek-V3.1 extracted intensifiers under formal or friendly, zero-shot or few-shot prompts. Twenty-four videos promoted gomutra and six debunked it. Promoting videos contained roughly 4.4 traditional metaphors and 2.8 scientific terms per 100 words, compared with 1.0 and 1.5 in debunking videos.
The models were useful but unstable. GPT-4o’s F1 was 52% for scientific terms and 59% for traditional terms. Cross-model intensifier agreement ranged from κ=−0.256 to 0.214, and changing prompt tone did not produce consistent gains. Crucially, competent debunkers also used culturally resonant language, meaning a detector could mistake counter-speech for promotion. The authors therefore argue that cultural competence cannot be added through surface prompt engineering alone. Reliable systems need local expertise, balanced multilingual data, medical evidence, subgroup audits, and human oversight.
Deep Analysis
Background
India has about 491 million social-media users, while YouTube grew by roughly 29 million users between January 2024 and January 2025. Prior health-misinformation work emphasizes COVID-19, WhatsApp diffusion, or English evaluation. Of 519 LLM health evaluations from 2022–2024, about 95% were in English. Gomutra discourse lacked an annotated dataset and systematic linguistic analysis.
Core Problem
The task is not merely checking whether “cow urine cures constipation” conflicts with medical evidence. Systems must distinguish promoters from debunkers who may both invoke religion, tradition, science, authority, and personal testimony. Video tone and gendered rhetoric add implicit context that surface lexical cues cannot reliably represent.
Innovation
- ��A first preliminary multilingual gomutra transcript corpus.
- ��Density analysis of traditional metaphors and scientific terms across stances.
- ��Intensifiers as an interpretable persuasion marker.
- ��A factorial comparison of three LLMs, formal/friendly personas, and zero-/few-shot prompts.
- ��Cohen’s κ analysis exposing reliability and gender-fairness risks rather than reporting only model outputs.
Methodology
- ��Collection: YouTube Data API v3 queries used cow urine, cowurine, gomutra, #gomata, and #cowurine.
- ��Processing: yt-dlp extracted audio; Whisper large transcribed it, with 7.04% WER on 16% of videos; GPT-4 translated non-English transcripts.
- ��Annotation: one author labeled promotion/debunking and speaker gender, producing 33 speakers.
- ��Feature extraction: GPT-4o identified traditional metaphors and pseudo-scientific terms, normalized per 100 words, and was compared with native-speaker annotations.
- ��Prompt study: GPT-4o-mini, Gemini 2.5 Pro, and DeepSeek-V3.1 extracted unique intensifiers under four conditions; precision, recall, F1, and Cohen’s κ quantified performance and agreement.
Experiments
The corpus included 30 videos: 24 promoting and 6 debunking, with 20 male and 13 female speakers. A random one-third subset received manual term annotation by a Hindi native speaker. Intensifier extraction compared formal versus friendly prompts and zero-shot versus few-shot settings. Agreement used binary presence matrices and Cohen’s κ. The paper did not include a conventional classifier baseline or an independent medical fact-checking benchmark.
Results
Promotional content had approximately 4.4 traditional metaphors and 2.8 scientific terms per 100 words, versus 1.0 and 1.5 in debunking content; reported tests included t=2.61 and t=2.63. Men used absolute-certainty amplifiers at 5.64 per 1,000 words versus 0.67 for women, and first-person “I” at 10.27 versus 2.61. Prompt changes altered outputs without a stable quality improvement.
Applications
Platforms can use LLMs for triage, transcript organization, and candidate-feature discovery, but not autonomous removal. Public-health agencies can commission culturally legible rebuttals while separating tradition, evidence, and unsupported treatment claims. Any deployment should audit language, gender, stance, and false-positive rates with local experts.
Limitations & Outlook
The small, imbalanced corpus, single annotator, translation pipeline, restricted languages, and API ranking limit generalization. Model outputs vary with versions and prompts, and intensifier frequency is not equivalent to medical falsity. Future systems need larger multi-annotator datasets, evidence retrieval, multimodal video analysis, and locally grounded language models.
Plain Language Accessible to non-experts
Imagine a very clever newspaper editor who grew up in another country. The editor can spot words that sound scientific, such as “antioxidant” or “detox,” and can count strong words like “permanent.” But it may not understand what a sacred cultural expression means to local viewers. The study showed this editor 30 YouTube videos: some claimed cow urine could cure constipation, obesity, or cancer, while others challenged those claims. To persuade believers, the debunkers also used familiar traditional language—like a teacher explaining a difficult lesson through a student’s favorite game.
The editor’s judgments changed when the request became formal or friendly, and three editors often disagreed. Therefore, traditional language cannot automatically mean misinformation, and scientific vocabulary cannot automatically mean truth. The safest newsroom would let the AI organize material, then ask local language experts and medical professionals to review it. AI can speed up the work, but it should not be the final judge of culturally sensitive health claims.
ELI14 Explained like you're 14
Picture yourself moderating a school group chat about health. Someone posts that cow urine can permanently fix constipation, help weight loss, and detox the body, adding medical-sounding words and a supposed patent. Another person wants to debunk the post, but uses familiar traditional ideas so classmates will listen. Would a robot mistake the debunker for the original spreader? That is the paper’s big question!
The researchers collected 30 Indian YouTube videos. Whisper turned speech into text with about 7.04% word error, and GPT-4 translated some transcripts. GPT-4o searched for traditional and scientific language, while three models searched for words like “absolute” and “permanent.” Twenty-four videos promoted cow urine and six opposed it. Promoting videos used about 4.4 traditional metaphors per 100 words; debunking videos used about 1.0.
The robot was not consistently reliable. GPT-4o’s F1 score was 59% for traditional terms and 52% for scientific terms. The models also disagreed a lot, and making the prompt friendlier did not magically fix the problem.
So AI is not useless—it is more like a fast gaming teammate who can collect clues but does not understand the whole strategy. Local experts, doctors, and multilingual reviewers still need to check the final call. Pretty important, right?
Glossary
Large Language Model (LLM)
A model trained on large text collections that predicts or analyzes language from context. It recognizes patterns but does not automatically possess cultural knowledge or medical judgment.
The paper evaluates GPT-4o, Gemini 2.5 Pro, and DeepSeek-V3.1.
Culturally embedded misinformation
Health misinformation whose credibility depends on religion, tradition, identity, or local authority. Its persuasive mechanism cannot be captured by sentence-level falsity alone.
Gomutra discourse is the case study.
Traditional metaphor
Language linking a claim to heritage, Ayurveda, sacred imagery, or traditional healing. It may support promotion or culturally effective debunking.
GPT-4o counted it per 100 words.
Intensifier
A word or phrase that increases force, certainty, or degree, such as “permanent” or “absolutely.” Intensifiers can signal exaggerated health claims but are not proof of falsehood.
Three models extracted unique intensifiers under four prompt conditions.
Word Error Rate (WER)
A transcription metric based on substitutions, deletions, and insertions relative to a reference transcript. Lower WER indicates closer agreement with human transcription.
Whisper achieved an average sampled WER of 7.04%.
Cohen’s kappa
An agreement statistic that corrects for chance agreement. Values near zero indicate little agreement; negative values indicate agreement below chance expectation.
It measured cross-model and prompt-condition agreement.
Open Questions Unanswered questions from this research
- 1 It remains unknown whether a larger, balanced corpus annotated by multiple local experts would substantially improve promotion–debunking discrimination. The study cannot estimate population-level performance from 30 API-ranked videos.
- 2 Traditional-language density is not equivalent to medical harm. Evidence retrieval, audience experiments, and longitudinal diffusion data are needed to connect rhetorical features with belief change and health behavior.
Applications
Immediate Applications
Human-in-the-loop moderation
A platform can use LLMs to transcribe videos and flag traditional metaphors or absolute claims, then route cases to Indian-language and medical reviewers. This reduces search effort while preventing autonomous takedown decisions and culturally driven false positives.
Localized public-health rebuttals
Health agencies can use culturally familiar framing to make corrections more persuasive, while explicitly separating religious tradition, clinical evidence, and unsupported treatment promises. Model-generated drafts should undergo expert and community review.
Long-term Vision
Culturally fair health-information infrastructure
Build multilingual, regionally diverse, gender-aware datasets linked to medical evidence graphs and auditable multimodal models. Such infrastructure could support cross-cultural evaluation and safer moderation, but requires sustained expert funding, governance, and community trust.
Abstract
Social media platforms have become primary channels for health information in the Global South. Using gomutra (cow urine) discourse on YouTube in India as a case study, we present a post-facto Large Language Model (LLM)-assisted discourse analysis of 30 multilingual transcripts showing that promotional content blends sacred traditional language with pseudo-scientific claims in ways that sophisticated debunking content itself mirrors, creating a rhetorical register that LLMs, trained predominantly on Western corpora, are systematically ill-equipped to analyse. Varying prompt tone across three LLMs (GPT-4o, Gemini 2.5 Pro, DeepSeek-V3.1), we find that culturally embedded health misinformation does not look like ordinary misinformation, and this cultural obfuscation extends to gendered rhetoric and prompt design, compounding analytical unreliability. Our findings argue that cultural competency in LLM-assisted discourse analysis cannot be retrofitted through prompt engineering alone.