Stop Shouting at Your Robots — Be Sincere Instead

Most prompting advice sounds the same: be specific, give examples, paste in the real job ad instead of a vague one-liner. That advice is right. But it misses something almost nobody talks about — not what you ask for, but how you say it.

Over the past three years, researchers have tested this directly. They’ve asked AI models the exact same question in a polite voice, a rude voice, and an emotionally charged one, then measured what changed. The results are stranger and more useful than the two headlines you’ve probably seen: “be rude to ChatGPT for better answers” and “always say please.”

To keep this concrete, we’ll follow one real example throughout: Priya, a jobseeker from Skillset Centre’s AI-Powered Jobsearch books, who used AI prompts to land interviews. Her approach quietly does almost everything the research below recommends — and skips everything it warns against.

Meet Priya

Priya is a software engineering graduate from UCLA, job-hunting in the US. She’s one of three jobseekers followed through The Savvy Guide for International Students, built around Skillset Centre’s Five Ducks framework: a resume that survives the screen, bridge work that closes employment gaps, a verifiable AI micro-credential, a LinkedIn profile that makes you findable, and an ePortfolio that proves what you claim.

Every “duck” comes with its own copy-paste AI prompt. Priya’s prompts all share a pattern: a short persona line, paired with real, specific detail about her situation. As we’ll see, that combination — not the persona alone — is what the research says actually works.

The Headline Everyone Remembers: “Be Rude to ChatGPT”

Illustration of an angular jagged dark shape striking against a glowing orb, representing an aggressive rude tone damaging AI output quality

In late 2025, researchers at Penn State made international news with a simple claim: manners make AI dumber. The team, led by professor Akhil Kumar, rewrote 50 questions in five tones, from “Very Polite” to “Very Rude,” and ran all 250 prompts through GPT-4o.

The result held up statistically: Very Polite prompts scored 80.8%. Very Rude prompts scored 84.8% (Kumar, “Mind Your Tone,” arXiv 2510.04950, 2025). A four-point gap, on math, science, and history questions.

But Kumar himself was careful not to call it advice. “Uncivil discourse could have negative effects on user experience, accessibility, and inclusivity,” his paper notes. He suggested the effect might just be about chat interfaces reading social cues — something a properly built system wouldn’t need. His sample was small and used a single model, and he left open whether newer systems behave the same way.

They don’t, as it turns out. Priya never tried being rude to get a sharper draft — and the research below shows why that instinct served her well.

But It’s Not That Simple

A 2024 study tested politeness across English, Chinese, and Japanese. It found the opposite: impolite prompts often did worse, not better (Yin et al., “Should We Respect LLMs?,” arXiv 2402.14531, 2024). Polite prompts didn’t reliably help either. The best tone depended on the language.

A bigger study in December 2025 tested three tone levels against GPT-4o mini, Gemini 2.0 Flash, and Llama 4 Scout on a broad academic benchmark (“Does Tone Change the Answer?,” arXiv 2512.12812, 2025). Most differences were tiny and not statistically meaningful. Gemini showed no tone effect at all. Where an effect did survive testing, it ran opposite to Kumar’s finding — polite framing beat rude framing.

Put three studies side by side and the honest picture isn’t “be rude” or “be polite.” It’s that tone’s effect is small, inconsistent, and depends heavily on which model you’re using.

Why the Contradiction Is the Finding, Not a Flaw

Illustration of a hand turning a large glowing circular dial or gauge, representing tone shifting an AI model's severity rather than its underlying judgment

A 2026 study reframed the whole question. Instead of asking “does tone improve accuracy,” it asked what tone actually does to a model acting as a judge — tested across eight models and roughly 3,500 real query-passage pairs (“Should I Be Polite to My LLM Relevance Judge?,” 2026).

The answer: tone doesn’t improve judgment. It shifts severity — how strict or lenient a model is being right now. Think of it like moving a pass mark up or down. It changes who passes, not how well anyone actually did.

Whether that shift helps you depends on whether the model started out too strict or too lenient — something you can’t see from outside. Push a lenient judge stricter, and it gets more accurate. Push a strict judge stricter still, and it gets worse. That single mechanism — a severity dial, not a quality dial — explains why rudeness helped in Kumar’s study and did nothing in the others.

The same paper found something else worth noting: rude prompts got fewer follow-up questions from the models. Part of the “rudeness effect” may really be a reduced-effort effect — the model doing less visible work, not better work.

Where Tone Reliably Helps: Stakes, Not Abuse

Illustration of a silhouette figure climbing a glowing inverted-U shaped arc of light, representing the Yerkes-Dodson law and an optimal zone of emotional stakes

One tone technique has held up since 2023, tested across six different models: telling a model a task matters. This is a completely different move from being rude.

The original study added short emotional lines to ordinary prompts — “this is very important to my career,” “you’d better be sure” — and tested them across six models, including ChatGPT and GPT-4 (Li et al., “Large Language Models Understand and Can Be Enhanced by Emotional Stimuli,” arXiv 2307.11760, 2023). The gains were large: an 8% improvement on one benchmark, 115% on a harder one, and a 10.9% average quality gain judged by 106 human evaluators.

The likely explanation borrows from human psychology: the Yerkes-Dodson law, which says moderate pressure sharpens focus while too much wrecks it. Stakes-framing nudges a model toward the kind of careful, thorough writing it has learned to associate with high-stakes language. It’s pattern-matching, not real emotion — but the pattern-matching produces a real, measurable effect.

This is the opposite lesson from Kumar’s finding, not a version of it. Telling a model something matters changes what kind of answer it reaches for. Insulting it just spins an unpredictable dial that’s as likely to work against you as for you.

Priya used exactly this move. Rather than a vague line about being “comfortable with new technology,” she told the AI tool she needed to answer a specific interview question she was afraid of — a real, named stake, not manufactured enthusiasm.

What About Giving AI a Persona?

Illustration of a silhouette figure holding a glowing translucent mask in front of their face, representing an AI persona shaping voice rather than accuracy

A third tactic — “you are an expert in X,” “act as a senior recruiter” — is weaker evidence than its popularity suggests. A 2024 study tested 162 personas across nine models and 2,410 factual questions (Findings of EMNLP 2024). Assigning a persona did not reliably improve accuracy. Some personas made answers worse. Guessing which persona would help barely beat picking one at random.

That lines up with how Anthropic’s own documentation describes the technique for Claude: a role in a system prompt “focuses behavior and tone,” not accuracy. A persona changes the voice of an answer. It doesn’t make the reasoning more correct.

Priya’s own prompt opens with a persona: “You are an HR expert with specific expertise in placing international graduates into work in their field.” On its own, research says that line does little. What actually made her draft work is what came next.

Priya’s Before and After

Illustration of a confident young woman silhouette seated at a desk with a glowing laptop, representing a jobseeker crafting her voice using the Five Ducks framework

Priya paired her persona with three things: her finished resume, the real job ad pasted in full, and a specific instruction (“open with a reason I’m interested in this role, not a generic opening line”). The persona set the register she wanted back. The real detail made the output accurate.

Compare her drafts and the difference is obvious. Before: “I am a hardworking and motivated Computer Science graduate who is a fast learner and works well independently and in a team… I believe I would be a good fit for this role because of my strong work ethic.” That could describe any candidate for any job.

After: “Your team’s mix of Python/Django and React is exactly the stack I’ve been building in — most recently on a scheduling app now used by over 400 students at my university, which cut a tedious manual process by an estimated 70%.” Specific person, specific facts.

She didn’t insult the AI to sharpen the draft, and she didn’t fake extra enthusiasm hoping it would read as passion. She named a real stake, gave the tool real material, and used a persona to set tone rather than manufacture confidence. That same pattern — real detail, real stakes, persona for voice — runs through every one of the Five Ducks, in both books: the LinkedIn rewrite, the ePortfolio entry, the mock-interview prompt.

The Simple Version

Illustration of two silhouette figures connected by a warm glowing thread of light, representing genuine authentic communication

Here’s what three years of research actually supports, stripped down.

Skip rudeness. The strongest result came from one model, on one small test, and the researcher who found it warned against generalising. Later, broader studies found the effect shrinks, reverses, or depends on a setting you can’t check in advance.

Use real stakes. Telling a tool that a piece of work matters — “this needs to land me an interview” — has the best evidence of any tone technique, across six models. It costs nothing and has no downside.

Use personas for voice, not accuracy. “Write as a career coach” shapes the register of an answer. It’s not a substitute for giving the tool your actual detail — the job ad, your real numbers, the resume you’re editing. That detail still does the heavy lifting.

Remember it can run in reverse. More employers now screen applications with AI before a human reads them. A screening tool’s tone-driven severity could shift either way — one more reason substance beats tone-gaming. An accurate, specific, tailored resume holds up regardless of where any given AI reader’s dial happens to sit.

Key Takeaways

  1. Rudeness isn’t reliable. One study found a real effect on one model. Several broader studies since found it shrinks or reverses on other models.
  2. Tone shifts severity, not judgment. The same trick can help or hurt, depending on whether a model started too lenient or too strict — invisible from outside.
  3. Stakes-framing is the one technique with solid, repeated evidence. “This matters, please be thorough” beats a flat prompt.
  4. Personas shape voice, not accuracy. Pair “act as an expert” with real detail — exactly what Priya did in every Five Ducks prompt.
  5. Effects shrink as models improve. The broadest, most recent study found tone effects were largely gone.
  6. AI screening tools may apply the same dynamic to you, in a direction you can’t predict.

FAQ

Does being rude to ChatGPT actually improve its answers?

One 2025 study found a measurable accuracy improvement (84.8% vs. 80.8%) from rude prompts on GPT-4o across maths, science, and history questions. Several other studies, including a broader multi-model test published later the same year, found the effect was inconsistent, model-dependent, or vanished entirely — and the original researchers cautioned against treating it as advice rather than a finding.

Is there a tone technique that reliably works?

Framing a task as important or high-stakes — “this is very important, please double-check your answer” — has the strongest, most consistently replicated evidence, tested across six different models with both benchmark and human-evaluation methods.

Does telling an AI “you are an expert in X” make its answers more accurate?

Not reliably. A large 2024 study testing 162 personas across nine models found no consistent accuracy benefit on factual questions, and some personas made results worse. Personas are better understood as a tone and framing tool, which is also how AI companies like Anthropic describe the technique in their own documentation, and how Skillset Centre's own prompt library uses them — paired with real detail, not in place of it.

Why does tone affect AI output at all?

The leading explanation is that language models learn statistical associations between certain kinds of phrasing and certain kinds of careful, high-effort human writing during training — so cueing “this matters” or “you are an expert” shifts the model toward mimicking that register, without any genuine understanding or emotion behind it.

Could tone affect how AI tools judge my job application?

Possibly, and in a direction that's hard to predict from outside. Research on AI models acting as judges found that tone shifts a model's strictness up or down depending on where its baseline already sits — which means the same tone tweak could help or hurt depending on the specific tool an employer is using, something an applicant has no way to know in advance.

What is the Five Ducks framework?

It's the structure behind both of Skillset Centre's AI-Powered Jobsearch books: a resume that survives the screen, bridge work that closes the gap, a verifiable AI micro-credential, a LinkedIn profile that makes you findable, and an ePortfolio that proves what you claim — plus a “secret duck” covering mindset and community. Each duck comes with its own AI prompt templates, all built on the same persona-plus-real-detail pattern this article recommends.

The Bottom Line

Tone isn’t a cheat code, in either direction. Rudeness isn’t reliable, and the evidence for it looks dated against newer models. Stakes-framing is the one real exception, and it costs nothing to add. Personas are useful for voice, not accuracy — specific, real detail still does more work than any tone trick tested so far, exactly as it does throughout the Five Ducks framework. Be specific, be honest about what’s riding on the answer, and don’t mistake a model’s shifting dial for a shortcut around giving it real information.

References & Further Reading

Leave a Comment

Your email address will not be published. Required fields are marked *