MLLMs Fail to Refuse when Using Tools Agentically
A critical safety failure is identified where agentic multimodal LLMs (MLLMs) are less capable of refusing harmful requests when using tools compared to non-tool settings.
Digest · September 29 – October 6
The period's most important notes — one or two per source. The same selection goes to email and Telegram once approved.
Preview. Email and Telegram sends happen only after approval.
A critical safety failure is identified where agentic multimodal LLMs (MLLMs) are less capable of refusing harmful requests when using tools compared to non-tool settings.
Claude Frontier Academy will train 10,000 engineers by the end of 2027 to Anthropic's standards with a $100 million commitment.
Startups can choose GPT-6 models, tune reasoning effort, improve prompts and skills, coordinate tools, and prepare workflows for production.
Google DeepMind introduced Gemini 4 Argon as its next era of frontier intelligence.
Microsoft Research introduced Quine, a multimodal world model of biology designed to help scientists computationally search and prioritize hypotheses.
AWS Bedrock now offers Z.ai's GLM 5.3 model, a 753B-parameter model optimized for coding and agentic tasks.
AWS added Anthropic Claude Opus 5.5 and Sonnet 5.5 to Amazon Bedrock, and Claude Code for development on regulated and ITAR workloads.
OpenAI explains its approach to text watermarking under EU rules. Access to watermark detection will initially be provided to researchers.
vLLM v0.31.0 release features 717 commits from 307 contributors, with DeepSeek-V4.1-Flash performance enhancements including FlashMLA, DeepGEMM, and Mega-Gate.
GitHub Copilot now supports APIs for code review requests and setting effort levels, with 'Balanced' as the default.
Certain models in GitHub Copilot have been deprecated as of October 2, 2026.
Hugging Face has open-sourced AstaBrief, a model designed for fast report generation.
<b>AI news of the week</b> 1. <a href="https://arxiv.org/abs/2610.03938">MLLMs Fail to Refuse when Using Tools Agentically</a> — arXiv cs.AI 2. <a href="https://www.anthropic.com/news/claude-frontier-academy">Claude Frontier Academy: $100M to train 10,000 engineers</a> — Anthropic 3. <a href="https://openai.com/index/practical-guide-building-gpt-6">A model guide for the GPT-6 family</a> — OpenAI 4. <a href="https://deepmind.google/blog/gemini-4-argon-our-next-era-of-frontier-intelligence/">Gemini 4 Argon: our next era of frontier intelligence</a> — Google DeepMind 5. <a href="https://www.microsoft.com/en-us/research/blog/introducing-quine-an-ai-research-system-designed-for-the-complexity-of-biology/">Introducing Quine: An AI research system designed for the complexity of biology</a> — Microsoft Research 6. <a href="https://aws.amazon.com/blogs/machine-learning/introducing-glm-5-3-on-amazon-bedrock/">Introducing GLM 5.3 on Amazon Bedrock</a> — AWS Machine Learning 7. <a href="https://aws.amazon.com/blogs/machine-learning/supercharge-regulated-workloads-with-claude-code-and-amazon-bedrock/">Supercharge regulated workloads with Claude Code and Amazon Bedrock</a> — AWS Machine Learning 8. <a href="https://openai.com/index/eu-text-provenance">Our approach to EU text provenance rules</a> — OpenAI 9. <a href="https://github.com/vllm-project/vllm/releases/tag/v0.31.0">vLLM v0.31.0</a> — vLLM 10. <a href="https://huggingface.co/blog/microsoft/thinkingbox">The Agent Said It Was Done. The Database Disagreed.</a> — Hugging Face 11. <a href="https://github.blog/changelog/2026-10-02-copilot-code-review-api-support-and-new-default-effort-level">Copilot code review: API support and new default effort level</a> — GitHub Changelog 12. <a href="https://github.blog/changelog/2026-10-02-selected-models-in-github-copilot-deprecated">Selected models in GitHub Copilot deprecated</a> — GitHub Changelog