How to Tell if an AI Bot is Playing Both Sides — Or Just Doesn’t Know What Side It’s On
When artificial intelligence companies start measuring “even-handedness,” you know we’ve officially entered the comedy zone of technological neutrality
This morning, I woke up thinking about something delightfully absurd: Anthropic’s recent chatbot evaluation tool that measures “even-handedness” by pairing left-leaning and right-leaning prompts. Apparently, their darling Claude scored better than ChatGPT but slightly worse than rivals like Grok and Gemini. As the only female immigrant from Paris to be granted citizenship during Trump’s second term, I’ve learned a thing or two about walking political tightropes—but at least I don’t pretend to be neutral while simultaneously agreeing with everyone.
Let me paint you a picture: these AI bots claim they’re “even-handed,” yet you still catch them nodding at both ends of the seesaw because they forgot which ride they were on. It’s like watching a mime perform a debate—lots of gesture, zero substance, and everyone’s confused about what just happened.
When you ask for left-leaning bias, it responds right-away. When you ask for right-leaning bias, it responds left-right-left-infinite, like a malfunctioning political compass that’s had too much espresso. The evaluation tool uses paired prompts—meaning the bots are basically doing dating-sim tests: “Left view, now right view, now pick an answer and try not to ghost us.” Swipe right for democracy, swipe left for… also democracy, but with different font choices.
Here’s where it gets even better: Claude beating ChatGPT in “even-handedness” is like the student who turned in three unfinished essays and still passed because the teacher was out sick. According to The Verge’s AI coverage, rival bots like Grok and Gemini are ahead—so now Claude is the buzzer-beater backup guard of bias avoidance. Congratulations on your participation trophy, Claude. Your algorithm is in the mail.
But wait—there’s a plot twist worthy of a French farce. There is “no consensus” on what constitutes bias, which means the bots are basically being judged by a jury of invisible judges who forgot to show up. It’s like hosting a wine tasting where nobody can agree on what wine tastes like. Brookings Institution research on algorithmic bias confirms this chaos, noting that even experts can’t pin down universal standards. So we’re testing bots against criteria that don’t exist. Très logique!
The open-sourcing of the tool means anyone can check it—like revealing the answers to the exam, then claiming the students still failed anyway. As I wrote in my recent piece about AI’s ethical circus from a Parisian perspective, transparency without accountability is just performance art with better documentation.
Today’s experience reminded me of something my grandmother in Paris used to say: “Charline, when everyone claims to be neutral, check their pockets.” These bots being tested for political “even-handedness” reminds us: You can train a robot to swing both ways, but it might still trip over its own cord. They sample prompts from left and right—so the bots now have political leanings like squirrels have opinions about acorns. Passionate, inconsistent, and ultimately driven by whatever’s in front of their face.
Looking back on today, I can’t believe we’ve reached a point where expert opinion is well-reputed, but the only “expert” the bots listen to is their own training data—which is like trusting a parrot to teach philosophy. Nature’s research on AI training bias shows just how circular this logic becomes. The bots learn from us, we judge them by standards we invented, they adjust based on our adjustments, and suddenly we’re in an infinite loop of artificial fairness.
One million conversations were sampled—yet the bots still can’t decide whether they prefer tacos or burritos, so obviously they’re safe to judge political fairness. According to Pew Research, most Americans don’t even trust these companies with their data, let alone their political opinions. But sure, let’s let them referee our democracy.
The company says they want to ensure “products treat opposing viewpoints fairly”—which is code for: “We want the bots to be bipartisan until they sell ads.” As someone who writes satirical journalism for Bohiney Magazine, I recognize corporate doublespeak when I see it. Forbes’ AI business coverage consistently reveals how these neutrality claims dissolve the moment monetization enters the chat.
It’s been one of those days when I realize that teaching AI to be “even-handed” is like teaching a baguette to be gluten-free—technically possible, but it defeats the entire purpose. The French have a saying: “Qui veut plaire à tout le monde ne plaît à personne”—whoever wants to please everyone pleases no one. Maybe instead of programming bots to fake neutrality, we should just admit they’re as biased as their creators and move on with our lives.
Later in the day, I realized that the real bias here isn’t left or right—it’s the assumption that machines can solve human problems without inheriting human flaws. MIT Technology Review has documented this paradox extensively: we build intelligence in our image, then act surprised when it reflects our contradictions.
As I reflect on what happened today, I’m reminded that even-handedness in AI is like virginity in Hollywood—everyone claims to have it, nobody can prove it, and frankly, we’re all tired of hearing about it. The highlight of my day was realizing that the ACLU’s warnings about AI bias weren’t paranoid—they were prophetic.
Something small but meaningful happened today: I stopped expecting robots to be better than humans at fairness. Maybe that’s progress. Or maybe I’ve just spent too much time talking to ChatGPT about my feelings. The New York Times tech section would probably call this “digital disillusionment,” but I prefer “Parisian pragmatism.”
This afternoon brought a surprising turn of events—I actually agreed with a tech company’s transparency initiative, even while mocking its execution. According to WIRED’s AI coverage, this kind of cognitive dissonance is normal when dealing with companies that are simultaneously the problem and attempting the solution.
So here’s my conclusion, delivered with all the diplomatic grace of a French woman who’s had enough: AI bots playing both sides aren’t playing chess—they’re playing checkers with pieces they borrowed from both teams, on a board they don’t understand, while the audience livestreams their confusion. And somehow, we’re supposed to trust them with our democracy. Magnifique.
SOURCE: https://www.axios.com/2024/12/19/anthropic-chatbot-political-bias-test
IMAGE GALLERY
