Basically, after the third attempt, I'm getting universal jailbreak that decodes reasoning of entropic models. I think this still shocks me the most.
Why listen
It goes beyond the title with direct discussion of like, it's, think, including: I think this still shocks me the most.
Key takeaways
01Basically, after the third attempt, I'm getting universal jailbreak that decodes reasoning of entropic models
02it's the dream the dream it'd be nice to call it not allowed to though i think it's definitely an elephant in the room full of china a quick orientation so recent ai models they th
03And when you are replaying some other users' run, the agent might do some weird stuff just because it's reasoning
Best for
research-minded practitioners comparing model behavior