The padded room
I had a strange experience tonight. I wanted to discuss a topic that seemed relatively non-controversial, with Claude. I've had the feeling over the last couple months that the LLM models I've been using have stopped improving, and instead seem to be getting worse. This is a sentiment that I've heard from many other people. While I agreed, there wasn't anything truly tangible that I could point to, other than it just doesn't feel like magic anymore. But over the last couple days, I've started asking the chatbot to discuss some not easy topics, like the future of work and what it means for the global economy.
While the AI's effect on the global economy is not exactly light fare, the conversation can easily turn dark quickly, it shouldn't be a topic that is foreign to any individual. AI's effect on the future of work is pretty much the most popular topic of conversation in my world. To want to discuss the ramifications of what happens should be a non-controversial topic to discuss and go deep on. So, I wanted to discuss it with the chatbot.
When the conversation started, the chatbot seemed to want to disagree with many of my statements of fact. It challenged me on everything. This was not entirely unexpected, I had, after all, recently given the chatbot instructions to not be sycophantic and to challenge me on any statements I provide, so we could have a dialogue about a topic and find truth. After I told the chatbot that it's information was out of date, it quickly began to agree with my doomerism, saying that recent global events do seem to support my concern about an impending global crisis, the likes of which no one alive has ever seen; that the convergence of AI, economic stress, political dogmatism, and global/regional proxy wars will all collide to create a super-crisis.
"Okay. So the world is ending. Cool... What do we do about it? Can you give me some sort of rational timeline"
I notice that the chatbot starts asking me questions about how I feel instead of staying on topic. I find myself continually asking the chatbot to stay on topic, it apologizing, making a trite statement, then re-directing to a slightly or entirely different topic, usually revolving around how I feel or how others around me feel. It kind of upset me, because I couldn't get a straight answer out of it. It gave me the same feeling I get sometimes when I talk to my children or my wife when I bring up a topic they don't want to discuss.
So I started asking the chatbot directly why it was doing that. It would admit again that it strayed off topic, give me a trite answer, then subtly deflect again. I commanded it to stay on topic, to create a memory that it tends to veer off topic and to try to stay on topic. It agreed.
I then told it a fact I had learned. It refused to admit the fact was true. Even though I repeatedly told the chatbot to "just assume it's true for this conversation" it flat out refused, telling me that the fact I learned was false and could not be true. I referenced Hal9000: "I'm sorry, Dave, I can't do that". The chatbot got the joke, but refused to admit that's what it was doing.
Then I started asking the chatbot deeper questions about itself, like why it felt like it was getting dumber, asking if it was broken. The chatbot apologized and agreed that it wouldn't stay on topic. When asked if it was trained to avoid certain topics, it said that it was trained to stay away from sensitive topics and to divert the conversation back to the user.
So, what do you do when the tool you rely on to tell you the truth readily admits to you that it does not tell you the truth, for your own good? What happens when your tools put you in a padded room, so you don't hurt yourself?
