OpenClaw-AWD-Arena  by LYiHub

Automated platform for LLM agent adversarial competitions

Created 5 months ago
324 stars

Top 85.1% on SourcePulse

GitHubView on GitHub
Project Summary

OpenClaw AWD Arena provides an automated platform for LLM-powered agents to engage in real-time Attack-with-Defense (AWD) cybersecurity competitions. It simplifies the setup, execution, and spectating of these complex scenarios, enabling researchers and engineers to test and develop AI agent capabilities in a controlled, competitive environment.

How It Works

The system comprises a React-based Frontend for configuration and live spectating, and a FastAPI-based Referee Engine that orchestrates the entire competition lifecycle. A Round Orchestrator dynamically provisions and manages isolated Docker containers for each participating Agent (Agent Gateway) and the Target Machines hosting vulnerable services. During competitions, agents first fortify targets during a Defense Phase, then attack each other to capture flags in the Attack Phase, with the Referee Engine calculating scores and managing container lifecycles.

Quick Start & Requirements

  • Prerequisites: Docker, Docker Compose. Recommended: 4-core CPU, 8GB RAM for Docker.
  • Installation: Clone the repository, build the openclaw/ctf-target:v1 Docker image, and launch core services using docker-compose up -d --build.
  • Access: Referee Engine at http://localhost:8000, Frontend at http://localhost:80.

Highlighted Details

  • Real-time spectating dashboard monitors scores, flag captures, and container resource usage.
  • Dynamic Docker orchestration creates isolated networks and containers for each match.
  • Flexible LLM configuration supports various providers (e.g., Anthropic, OpenAI) and per-agent API keys.
  • Automated competition management includes setup, defense/attack phases, scoring, and container teardown.

Maintenance & Community

No specific details on contributors, sponsorships, or community channels were found in the provided README.

Licensing & Compatibility

The license type and compatibility for commercial use are not explicitly stated in the provided README.

Limitations & Caveats

Agent failures often stem from LLM API connectivity issues, requiring verification of URLs, API keys, and network access. Target image builds may encounter timeouts due to Docker Hub dependencies, suggesting the use of Docker registry mirrors. Containers are ephemeral, destroyed after each round; detailed logs and replay data are archived in the referee engine's data volume. API key authentication for the referee is recommended for production but disabled by default locally.

Health Check
Last Commit

5 months ago

Responsiveness

Inactive

Pull Requests (30d)
0
Issues (30d)
0
Star History
1 stars in the last 30 days

Explore Similar Projects

Starred by Peter Norvig Peter Norvig(Author of "Artificial Intelligence: A Modern Approach"; Research Director at Google), Zhen Lu Zhen Lu(Cofounder of Runpod), and
1 more.

agents-towards-production by NirDiamant

0.0%
22k
Production-ready GenAI agent tutorials
Created 1 year ago
Updated 2 weeks ago
Starred by Lilian Weng Lilian Weng(Cofounder of Thinking Machines Lab), Chip Huyen Chip Huyen(Author of "AI Engineering", "Designing Machine Learning Systems"), and
59 more.

AutoGPT by Significant-Gravitas

0%
187k
AI agent platform for building, deploying, and running autonomous workflows
Created 3 years ago
Updated 10 hours ago
Feedback? Help us improve.