Articles, research reflections, and technical essays on computational humor, language model evaluation, and modern NLP.
How evaluating computational humor exposes critical limitations in using large language models as judges, drawing from insights built during the HumorRank project.
Exploring the methodologies behind HumorGen-7B: aligning LLMs to generate high-quality humor through multi-agent cognitive synergy and LoRA distillation.
Designing robust evaluation frameworks that measure nuance, context sensitivity, and intent alignment in interactive and conversational agents.
Get notified whenever I publish new essays, research notes, and commentary directly to your inbox.
Subscribe on Substack