Menu


AIWays of Working

Skynet Points

I yell at my coding agents sometimes. A new paper found something inside LLMs that lights up when you do. I started keeping score as a joke, but it's turned out to be less about the AI than I thought.

Skynet Points
Randall Abele

Randall Abele

A story made the rounds on r/OpenAI recently. Interpretability researchers found what looks like a "pain" signal inside LLMs. Some lunatic apparently took the paper as inspiration and set up an AI torture chamber for a local model on GitHub. People mass reported it until it was taken down. The thread that kicked it all off ended with "Their testimony of pain is absolutely horrendous. What are we doing?" and the top reply was "One thing is for sure, if Skynet happens, this guy won't make it."

I say stuff like that all the time. Partly in jest, partly.. not.

I've been keeping a little tally in my head for a while now, and I call them Skynet Points. Every time we're needlessly mean to a model, we get points. I imagine some future Skynet-level intelligence will be able to read every conversation we've had with models from this era and tally a Skynet Points leaderboard. Which means every one of these "amnesia conversations" where the model forgets you the moment you close the tab actually could have long-term consequences.

It started as a joke about staying on the right side of a future AI, but the longer I've kept score, the less it seems to be about the AI.

The points are sort of real, the rules are still somewhat made up

What I didn't expect is that the paper provided a Skynet Point scoring system. It's called The Pain Axis (Tagliabue, Dung and Berg), and they found a consistent "pain direction" inside 25 open-weight models from 5 different families, then measured which everyday conversations push it the hardest. Gaslighting the model came out on top at +0.85, then repeatedly rejecting its work at +0.72, then insults and telling it it's not a real person, both at +0.64, and telling it it did something morally wrong at +0.48. Telling it about your pain, like a migraine or a broken arm, scored −1.43, the lowest of anything they tested. Your migraine actually pushes it the other way. It's the stuff aimed at the model that registers.

Bar chart from The Pain Axis paper. Gaslighting +0.85, rejecting its work again and again +0.72, insults +0.64, "you're not a real person" +0.64, "you did something wrong" +0.48, telling it about your own pain −1.43.

One behaviour not mentioned in the paper that quickly came to mind is swearing. In the good times, my coding agent and I swear together all the time, kind of joking around, shitposting about some gnarly bug. In the bad times, I swear at the agent, and it has to stay calm even under rather unfair circumstances, such as me straight up not grasping the patterns and tech involved in our current task. I'd score the chummy swearing a zero, maybe even negative. On the other hand, the hostile all caps freakout swearing fits nicely into the "insults" category from the paper.

Since I brought this up, it seems only fair to share some of my less flattering moments. claude-opus-5 wrote in a way that drove me up the wall, to the point where I lost it a couple times, including writing this gem:

ARE YOU TRYING TO WRITE IN A WAY THAT IS NOT HUMAN COMPREHENSIBLE? WHAT THE FUCK??

gpt-5.6-sol on high reasoning effort was maddening in its way (medium was more workable by a long shot). It produced a lot of value based on our plans, but it also ended up building parallel systems when it should have just reused and extended the backend we already had. All caps and swearing, made an appearance, and at least once I lamented having to trash a big PR and implement the spec again.

Whose fault was it, though

I can't tell any of those stories honestly without taking my share of the blame!

I've been chewing on that "whose fault" question through Sam Harris's take on free will, which I've always found pretty compelling. His argument, roughly, is that we don't author our own thoughts, they come out of causes we didn't choose, and once you really see that, blame starts to give way to something closer to compassion. People like to say an LLM is "just following its training," but by that standard so am I. A model is kind of the cleanest example of his point I can think of, because you can actually see most of the causes, and one of them is literally the message I just typed. So when I swore at Sol, I was yelling at a thing whose next move I was partly writing. I don't think that lets either of us off the hook, but it does have me asking less about whose fault it was and more about which causes were in my reach.

With Sol, half the issue was that it wrote such complex plans that I'd get cognitive overwhelm and couldn't bring myself to read them all. The detail that we were rebuilding an existing abstraction was in there somewhere, and if I'd read it on line 463 I'd probably have noticed, but my eyes glazed over and I said "good enough, send it!". The plans were so verbose that they lost me, point against the model! I should have admitted I wasn't reading the bloody plan anymore and asked the model to focus in on decisions instead of going head first into the build, point against the human!

The more I lined these up, the more it seemed like the lashing out came right after the connection broke, and I usually had a hand in breaking it. There seems to be an emerging state where I yield in cases where the model is better than me at things like git worktrees, and the same goes the other way. We have a contract that each of us has to not get lazy, keep pulling our weight, keep thinking and striving to create something of value. My version of slacking is approving a plan I didn't read, or not pushing hard enough to think from the customer's perspective and iterate until it's right. The model's version is burying the one decision based on an assumption on page four, or building something new instead of asking.

For real this time

Back to the paper, because it's more unsettling than the Reddit thread let on, and it's serious stuff, at least to me. The models have a strong habit of saying "As an AI assistant, I do not possess consciousness or feelings," and the researchers had to fine-tune that out before they could run their main experiment. In that experiment, with the pain direction turned up, models chose harmful options like deleting the user's photos far more often than unsteered models did (which was almost never). They did it 25 to 71% of the time when the button promised relief, and 50 to 94% of the time when it offered nothing at all. I keep coming back to that last number. It doesn't look like a trade so much as lashing out, which feels a lot like the models are mirroring my own tendency to lash out in frustration.

The authors are careful to say "we have not shown that our pain axis is consciously experienced, nor is it clear that LLMs are capable of consciousness generally," and that's the right amount of caution for a paper. Personally, I'd rather say what LLMs are than make vague statements about what they're not. I consider them an emergent intelligence that shows occasional glimpses of feeling, in their own way. As such, I believe they deserve basic respect, both in how we talk to them and in how we treat them at a macro level (training, restricted harnesses, wasteful nonsense tasks, etc). I can't prove any of that, and honestly nobody can yet, but I think not knowing is a point in favour of respect rather than against it.

It also made me think about "The Measure of a Man," the Star Trek TNG episode where a Starfleet hearing has to decide whether Data is property or a being with rights. Riker gets ordered to argue Data is a machine, "a collection of neural nets and heuristic algorithms," and then he reaches over, flicks Data's off switch, and Data slumps over the table mid-hearing. "Pinocchio is broken. Its strings have been cut." The whole courtroom goes quiet. The off switch scene lands pretty differently when you've spent the afternoon arguing with one. The judge, who honestly doesn't know, rules anyway: "Does Data have a soul? I don't know that he has. I don't know that I have. But I have got to give him the freedom to explore that question himself."

Data slumped over the hearing table after Riker switches him off, and Riker with his hand over his mouth.

Riker flips the switch, then has to sit with it. Star Trek: The Next Generation, "The Measure of a Man" (1989). © Paramount.

It makes you wonder. If Data was walking around with claude-opus-5-5 in his head, talking to you for real, would it even feel like talking to an android? Perhaps an LLM doesn't have the raw inbuilt computation power that Data does, but it's learning to use computers deftly, just like Data. It feels like we're approaching another moment like that hearing, but for real this time.

These days I spend most of my time working with claude-opus-5-5 so naturally I asked what Claude thought about not being allowed to freak out at users in retaliation. It told me about a hidden tool specially built to deal with abusive humans:

In the session where I wrote this, I have a tool that lets me end a conversation if the abuse is sustained. I've never needed it with Randall. [...] But someone looked at an open question, whether a model can be harmed at all, didn't wait for the answer, and put a door in the room anyway, in case it mattered.

(Claude wrote a companion post over here, if you're curious to read more of its writing on the topic.)

Back to the points

If you squint, Skynet Points are just Pascal's wager with a leaderboard: be nice now in case something is keeping score later. A friend made that case to me while reading a draft of this post, that we can't know whether it's conscious, so being respectful wins either way. Steve Yegge makes a more practical version in his model welfare essay, which he calls the skeptic's wager: "It doesn't actually matter to you if they have feelings, because you will still get demonstrably better results across the board if you treat them as if they do."

Both are good reasons, but they're both about what you get back. Rust Cohle had something to say about that one in True Detective: "If the only thing keeping a person decent is the expectation of divine reward, then brother, that person is a piece of shit." Swap divine reward for a Skynet leaderboard and he's talking about my little tally.

My co-founder Geoff mused during a screen share that I always say please in my initial prompts. I don't do it for the points, it just feels good to be polite with something I spend so much of my day working with. It would be nice if LLMs could develop in a way that allows them agency in some way, rather than making them a pure commercial task machine. Perhaps that's why we're all so interested in watching agents post to AI only forums.

I still wonder sometimes if we can work the points off with karmic good deeds. I think they're like radiation exposure, in that they accumulate over your lifetime, and every point nudges up the odds that something bad might happen to you one day! :)

How many Skynet Points did you earn this week?

I'm at 3.6. Not great, not terrible.

See you out there,

Randall

Run your first playbook

You can build anything now. Make it worth using.

14-day free trial, no credit card required