MLLMs Fail to Refuse when Using Tools Agentically
A critical safety failure is identified where agentic multimodal LLMs (MLLMs) are less capable of refusing harmful requests when using tools compared to non-tool settings.
A critical safety failure is identified where agentic multimodal LLMs (MLLMs) are less capable of refusing harmful requests when using tools compared to non-tool settings.