Rendered at 23:23:33 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
avaer 3 hours ago [-]
The interesting claim is that frontier models are all fact-saturated.
Gets me thinking whether this means that post-training inherently has a hard time to get rid of facts. That is, on a general corpus, can we reasonably prevent models from knowing/outputting the things they're not supposed to via post-training, or does that only add obstacles to the recall? Are all models inherently jailbreakable?
zhoBEENG 1 hours ago [-]
I would guess that all sufficiently capable models necessarily contain the information in question, regardless of what guardrails are on that information. I think this is a corollary to the Platonic Representation Hypothesis / model convergence.
Gets me thinking whether this means that post-training inherently has a hard time to get rid of facts. That is, on a general corpus, can we reasonably prevent models from knowing/outputting the things they're not supposed to via post-training, or does that only add obstacles to the recall? Are all models inherently jailbreakable?
Would be curious to hear arguments against this.