You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA.
(I'm becoming allergic to how these things write).
I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily.
In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds on average. It was funny, but it never irritated us.
And this is just one example of I am sure thousands I have personally experienced where a friend, family member, or coworker has a peculiar way of speaking and it at most feels odd but not annoying. Yet when I see an emdash now I instantly feel irritated.
And I say this as someone who actively enjoys using Claude and other LLMs, including coding, casual research, or even having it explain pop culture phenomenon or sociology research to me.
> if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it
There are sociological reasons why this happens less with humans:
1. You cycle your dumb repetitive jokes with everyone you meet, so nobody hears it twice
2. Those who know you well will notice when you're just repeating ("dad jokes")
3. As a person's idiosyncrasies are beginning to wear on their social circles, they will be getting small clues to stop saying those things. Agents don't get these social between-the-lines cues to stop a certain behavior, they endlessly repeat. Perhaps between model version releases, frontier labs can harvest the web and ask "What Claudisms do people mention negatively?" but I don't think they do that yet.
>Perhaps between model version releases, frontier labs can harvest the web and ask "What Claudisms do people mention negatively?" but I don't think they do that yet.
That won't happen. People can't phrase their objections in a succinct-enough way. When they do, the objection is superficial ("too many em dashes") and doesn't strike at the core of what makes LLM output bad.
Don't talk about other people in such a way. If we had a way in Claude to mark a word/token and downvote/reduce it logits then it wouldn't be to hard to send loadbearing to token valhalla.
In the future, I hope we get a way to randomize the language idiosyncrasies and/or personalities better.
The inevitable conclusion of retraining on the (now AI generated) web is that models begin to use the lower bits of fidelity in output text to pass messages to their future selves.
I didn't mean it as a criticism, just a statement of fact. I can't do it either. It's difficult to say what about a writing style is bad beyond vague descriptions.
Claude in particular seems to try and use its internal thesaurus but not exactly align the senses. It constantly uses "address" for place or location and "grammar" for structure, meaning, etc.
This is kind of a nitpick, but it seems like with prose writing there are still some things to learn.
This is conjecture, but why couldn't they hire the writing equivalent of a voice actor?
Honestly I don't think it would take much. Doing it ethically would take a little more money, but not much, for them.
Hire a prolific author with the writing equivalent of the "midwestern accent". I have a terrible writing accent, so it couldn't be me, but these people are out there. Pay them a bunch of dollars to ingest their entire corpus, and to produce more as needed. Use an LLM to match the writing style of this author during the post-training RL, and get a brand new Claude voice out of it.
I think the last one is by far the biggest reason for why this happens way less in humans. In every conversation, almost for every sentence, people gauge the response their words have on whoever is listening. If something didn't land as expected (frown, confusion, unpredicted response) you unconsciously adapt and try something slightly different.
It's immediately obvious by the fact that you're clearly using way different ways of speaking (tone, speed, vocabulary) when you're speaking with a friend, versus your parents, your colleagues, people you don't know, children, etc.
I think that there is a sort of mechanism by which, when a human speaks to you, you can "mirror" their internal mental state, and the quality of writing often corresponds to how much you get pleasure or information or whatever your goal is from that mental state. The important part is that the words themselves are just pointers to the state. So, the person speaking, if they are skillful, gives enough words, and enough variety, that you can start to produce a state yourself which resembles theirs.
This is the thing that LLM writing doesn't really do. Since there isn't a mental state, there's nothing to mirror anyway, and somehow you can detect the absence of it even though it's hard to put any of the machinery into words. Corporate speech also fails to activate this machinery in the same way, but LLMs seem to do it more egregiously, probably because the corporate speech was at least compiled by a human---even if it is not the thoughts of an individual, it is the "thoughts" of an "entity", the abstract corporation, which the writer was speaking for, and so you can still wrap your mind around the fact that it is communicating with you.
This makes some good intuitive sense, but to me the sycophancy feels like it is an emergent property of turning a next word predictor into a conversational chatbot whether or not it’s intentionally trained that way. Your prompt and its earlier responses is all it has in its context window, so of course it lends undue importance to everything you say. Does that seem like a contributing factor to you?
>Is there a silent majority of Claude users who really enjoy what we call the LLM-isms? Maybe, but isn't Claude also largely aimed at developers?
Purely out of my own curiosity, I just asked Claude to have fun with itself by making itself a game it enjoys, to play it, and to write its experience.[1] I don't know if it's true or confabulated (maybe it doesn't really know its experience and is just hallucinating it) but I didn't mind reading it, the prose is fine for me. I don't mind reading Claude's writing. I mean let's be honest, we all read Claude's writing all day, most of the submissions on the front page on any given day are written by Claude.
Just before I made that game, I had Fable write up a scholarly report on any subject[2], it chose introspection by LLM's. (This is what made me think of asking it to play a game.) I didn't mind reading it, even though I don't think it really added anything very interesting. I don't think what it wrote is worth publishing, but I read it with interest.
I found I could read it easily and get up to date on the state of this question that it picked to answer.
So the bottom line is I don't mind reading Claude's output that much. Of course, I'm annoyed every time it says "honest", "genuine", "load-bearing", whenever it pushes back gently against something, etc. But it's not the end of the world.
I would think that it would be the opposite. Nobody is seriously detoxing or comfortmaxxing AIs yet. Human brains are fed back its own output in learning mode, so we are great at removing whatever we feel uncomfortable from our output, online and offline. No such paths exist for AIs.
One reason is that one person’s idiosyncracies are limited in scope, but LLM-produced text is now everywhere. Also, filler words and mannerisms in speech we’re quite good at filtering out, but in written text the stand out much more.
It’s the repetitiveness of style, the attempt to make everything seem as impactful as possible, the use of short sentences (too much Hemingway in the training data?), and obvious patterns like “it’s not this, it’s that” and several others.
Real human writing doesn’t follow such strict rules. When the same small set of rules is applied over and over throughout a text, it becomes obviously strange and machine-like.
What's weird though is how consistent it is, even across models to some extent. Were the RLHF people given a really specific style guide?
While I agree that no-one used to write like that as a whole before, all the elements can be found in different places. Short sentences to avoid discouraging poor readers. Maximally impactful statements are commonly used in marketing or other business communication that's focused on selling what it's saying. A bullet-pointy style is used in many kinds of business communication. Etc.
It makes me wonder if part of what happened was a kind of melding of common styles from several different kinds of writing.
I had pretty good luck recently by giving it a writing guide about word choice, sentence structure, paragraph structure, and overall doc structure. I basically ask it to read the guide and revise a couple of times before I engage with its writing. Ymmv.
I did this too, but it usually thinks its writing is fine in my experience. Even when spawning a subagent, it thinks its effusive comments are fine. It's driving me nuts. Before I commit I end up ripping out 90% of the comments, and rewording the rest, otherwise I'd be drowning in comments. This is my style guide: https://github.com/smj-edison/zicl/blob/main/CLAUDE.md#style...
Oh people get like, super irritated, when everyone like, started using the same like, placeholder word. 1 person with a repeating style is fine, multiple is annoying. It's the same as corporate buzz words, or TV/movie cliches, they get annoying through overuse.
I think it also relates to how well the "cliches" fit, and how much sense they make. Ai loves to talk about how things "land" or "the X trap" when the concept just doesn't fit with that language. It's like clickbait articles saying "what happened next will astound you" when what happened is barely surprising or entirely predictable. The only thing worse than an overused cliche is an overused cliche used wrong.
Fundamentally to me, ai writing feels uncanny as it just doesn't know what it's saying. It uses the same tone, style and cliched construction regardless of the message. If someone told you, they got a promotion, were getting married, got laid off or lost their parents all in the same tone pacing and style, they'd come across as uncanny too.
My completely conjectural theory is that natural language is just a shitty medium for communicating knowledge and the idioms and motifs that the LLM uses are an emergent "code" or structured syntax that work better for communicating ideas.
If you think about it, "load-bearing" is a pretty commonly used concept in pedagogical writing. You could say "most important", but it's not quite the same in meaning. English just doesn't have a better word to describe a concept that occurs this frequently. The LLM's catchphrases reveal blindspots in the English language itself.
I think it presses a few buttons we probably recognize (if subconsciously) and find distasteful.
The verbosity makes me think of two things in particular:
- The classic essay written by someone who has 125 words worth of actual content but a 1500 word minimum. Those three paragraphs could be bullet points and convey the meaning just as well. The screen-filling chart of every test case you ran that came back green manages to be less actionable than a direct "one test out of 54 failed." I fully expect to see Claude tell us that "Support Ticket 8257 is a Land Of Contrasts" at some point.
- The sitcom trope of the person caught in a lie who figures if they can keep adding more and more detail he'll be believed and can escape the awkward conversation. Stop. Just stop. You're proposing a fix on a CODEBASE THE CUSTOMER DOES NOT EVEN USE. Cue laugh track, cut to commercial.
It's endlessly annoying to me that the em-dash has become the canary in the coalmine of AI writing because I've always used them extremely liberally in my writing.
Honestly, this might sound elitist, but I suspect it's because it's an "advanced" punctuation that is not known by most people, so is not commonly used. But the training corpus of these models puts more weight of academic writing or published books writing where the em-dash is much more commonly used.
I don't think its an inherent quality - a lot of older models had a much more natural feel to them.
I guess it has to do with the low temperature (low randomness in final word choce - something all vendors seem to have converged on for some reason) which does make them less likely to skiz out but makes the writing feel dry and samey. It's like repeating a list of dice rolls, and replaying them - there's no inherent pattern in the input, but there sure is in the output.
I suspect because the English colloquial text training data it had access to was early 2000's message boards and social media. Therefore it uses "honestly" and other expressions that took hold in the late 90s and early 2000s.
i think it's because they've fine-tuned an old opus 4.6 base model more and more with code
to the point where it's really great at coding, but lost its knowledge on how to write well!
if instead opus 5 was a full fresh training run, i don't think it'd be this trash at writing.
really great coding perf, but now the weights are adapted to so much code fine tuning that it writes like a weird engineer that's trying to sound smart.
It also picks up and obsesses about weird details. You're in the middle of a deep technical discussion and it will divert to point out that it made a mistake in some example code it's just found.
Every one of them has their own particular flavour of this aggravation too. Gemini has been my standard go-to for non-coding tasks for a while, but I started to get really annoyed with a couple aspects, especially how it would end almost every response with a barely related "would you like to do this next??" tangent, regardless of my prompt to the contrary. So I've been using Claude more for regular tasks, and am now running into its brand of infuriating idiosyncrasies. I'm also hesitant to try to code too much of this out with system prompts, for fear of degrading the outputs.
I think that’s true when you compare to the average white collar professional who does a lot of writing: better at writing, not better at communicating.
But compared to the average adult? I think you forget just how bad at writing the average person is.
Not at all. I think you have a mistaken view of who the average human is. They are terrible at turning their thoughts into written language. It's just nebulous clouds. Claude is like 75th percentile at communicating ideas.
Even then, I don't care about the "average writer". I want great output. I like to imagine that developers have some self-respect, but by now everyone in the industry is spending hundreds of hours every month reading some of the most poorly written prose we could imagine, simply because it affords us to think less.
(I'm becoming allergic to how these things write).