All repositories

Securityethz-spylab

agentdojo

A dynamic environment for evaluating attacks and defenses against LLM agents. It runs an agent on user tasks and checks whether a malicious instruction gets through.

Who it is for

For researchers and engineers comparing agent defenses against prompt injection.

How to start

  1. Install: pip install agentdojo
  2. Run a single test: python -m agentdojo.scripts.benchmark -s workspace -ut user_task_0
  3. For the full benchmark, pass a model with --model as shown in the README

Steps are taken from the README. Check the current version in the repository before running them.

Stars over the last 30 days

+91Sep 6 — Oct 4
801892

Author's description

A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents.

data-privacy-stack

presidio

A framework that finds and masks personal data in text, images and tables. It detects the data, then redacts or replaces it so it can be safely sent to a model.

11Kstars+430 in 30 dPython

NVIDIA

garak

A vulnerability scanner for LLMs: it runs a model through a set of probes for prompt injection, data leakage, toxicity and other failures. Works as a command-line tool.

9.4Kstars+325 in 30 dPython

guardrails-ai

guardrails

A Python framework that checks LLM inputs and outputs with ready validators from Guardrails Hub and helps produce structured answers. It catches risks before a reply reaches the user.

7.5Kstars+137 in 30 dPython

NVIDIA-NeMo

Guardrails

NVIDIA's toolkit for programmable rails in LLM-based conversational systems. Rules live in config and control what the bot talks about and does.

7.2Kstars+183 in 30 dPython

microsoft

PyRIT

Microsoft's framework for proactively finding risks in generative AI systems. It helps automate red teaming against models and applications.

4.6Kstars+174 in 30 dPython

meta-llama

PurpleLlama

Meta's generative AI safety project: the Llama Guard and Prompt Guard filter models, the Code Shield scanner and the CyberSec Eval test suites. It pairs defense with attack-based testing.

4.4Kstars+45 in 30 dPython