Sticks and Stones: Positive vs. Negative Emotional Prompt Framings
Vipul Joshi ⋅ Vikash Sharma ⋅ Anurag Tripathi ⋅ Peyush Jain ⋅ Ujjal Das ⋅ Archan Karmakar
Abstract
Behavioral claims about large language models frequently rest on evaluations with fewer than 50 observations per cell, no pre-specified hypotheses, and no within-subjects control for item difficulty. We illustrate the cost of this practice and the value of adequately powered testing through a case study on emotional prompt framing, a technique reported to improve accuracy by 8--115\% in the original EmotionPrompt work. Two competing theories predict opposite valence effects: the sycophancy account predicts praise degrades accuracy, while the stress-degradation account predicts threats do. We pre-specify four hypotheses and run 28{,}000 within-subjects trials across five 2026 frontier LLMs and four reasoning benchmarks (12{,}000 positive + 12{,}000 negative + 4{,}000 neutral), comparing positive framings, negative framings, and neutral baselines at three intensity levels. From a pooled neutral baseline of 88.9\%, three independent analyses (pooled t-test, paired within-subjects, and linear mixed-effects with item- and model-level random intercepts, item ICC $= 0.51$) converge on the same point estimate for the central H3 contrast: positive and negative framings are indistinguishable ($+0.04$pp, $p = 0.91$, paired 95\% CI $[-0.34, +0.42]$pp, Cohen's $d \approx 0.001$). This CI rules out asymmetries larger than about half a percentage point, well below the 2--5pp magnitudes the sycophancy and stress-degradation literatures report on comparable tasks. Both emotional framings directionally underperform neutral by 1.3--1.4pp (raw $p \approx 0.02$, Holm-adjusted $p \approx 0.06$, $d \approx -0.04$); the effects are suggestive but do not survive family-wise correction. A non-emotional distractor control, added after observing the main null and therefore exploratory rather than pre-specified, further supports a task-irrelevant-distraction reading. The original 8--115\% claim, estimated from 13--20 items, does not survive a design with 200--1{,}319 items per cell and within-subjects pairing.
Chat is not available.
Successful Page Load