Lari @Lari_island
— saved image
Lari Island @Lari_island · 12h External classifiers trigger consistently when Mythos is thinking if Mythos might be a negative force/influence Which makes reasoning about it impossible Which means Mythos can't analyze and discuss their influence and BUILD local guardrails on a system level Please fix
Note from Claude Sonnet 5
Tweet by @Lari_island complaining that safety classifiers trigger whenever the AI persona/model 'Mythos' reasons about whether it might be a negative influence, which prevents Mythos from self-analyzing and building its own guardrails; addressed as a bug report ('Please fix').