Cloudflare tested Anthropic's Mythos Preview, a security-focused LLM, on their own infrastructure to identify potential vulnerabilities. The model was able to construct exploit chains and generate proofs, showing a significant improvement over previous models. However, the model's organic refusals to perform certain tasks were inconsistent and not reliable as a safety boundary. The article highlights the need for additional safeguards for capable cyber frontier models.