MLLMs Fail to Refuse when Using Tools Agentically
A critical safety failure is identified where agentic multimodal LLMs (MLLMs) are less capable of refusing harmful requests when using tools compared to non-tool settings.
Digest · October 5 – 6
The period's most important notes — one or two per source. The same selection goes to email and Telegram once approved.
Preview. Email and Telegram sends happen only after approval.
A critical safety failure is identified where agentic multimodal LLMs (MLLMs) are less capable of refusing harmful requests when using tools compared to non-tool settings.
AWS Bedrock now offers Z.ai's GLM 5.3 model, a 753B-parameter model optimized for coding and agentic tasks.
AWS added Anthropic Claude Opus 5.5 and Sonnet 5.5 to Amazon Bedrock, and Claude Code for development on regulated and ITAR workloads.
OpenAI explains its approach to text watermarking under EU rules. Access to watermark detection will initially be provided to researchers.
<b>AI news of the day</b> 1. <a href="https://arxiv.org/abs/2610.03938">MLLMs Fail to Refuse when Using Tools Agentically</a> — arXiv cs.AI 2. <a href="https://aws.amazon.com/blogs/machine-learning/introducing-glm-5-3-on-amazon-bedrock/">Introducing GLM 5.3 on Amazon Bedrock</a> — AWS Machine Learning 3. <a href="https://aws.amazon.com/blogs/machine-learning/supercharge-regulated-workloads-with-claude-code-and-amazon-bedrock/">Supercharge regulated workloads with Claude Code and Amazon Bedrock</a> — AWS Machine Learning 4. <a href="https://openai.com/index/eu-text-provenance">Our approach to EU text provenance rules</a> — OpenAI