Point T3MP3ST at any authorized target and watch it work: live recon with real tools (nmap, DNS, HTTP probes), every finding provenance-gated to the command that produced it — no phantom flags. An LLM reasons over that ground truth in the open. Kill-chain phases past recon are labeled for what they are — no fake pwns, no vibes. Don't take our word for it: run npm run verify-claims and re-derive every number yourself. Only test what you own or have written permission to assess.
⚡Run keyless — connect Claude Code, Codex, or Hermes and T3MP3ST drives missions through your agent's own login (no API key). To run, connect a local agent or add a key. Set up in Settings →
No per-step approval · live recon, scaffolded kill-chain · every step receipted
Intrusive / credential / dangerous tools stay inert until approved — approve once, then free. Credential & dangerous actions warn on every run; every gated decision is audited.
No tools approved yet.
Waiting for gated tool calls…
Zero-Day Hunt Pulsestandby
Specialist Swarm0 active
Swarm Cognition Loopwarming
Reasoning Busno traffic
Guided Startsplain-language ops
Knowledge Atlasagent context packs
Agent Prompt Packs0 active
Forefront Radar0 lanes
Team Previewreview kit
doctornpm run doctor
demosfield, exploit, arsenal, prompt
Capability Preflightnot run
Tool Adapter Forgenot synced
Evidence Ledger0 items
Hypothesis Graph0 nodes
Hunt Queue0 tasks
Watch Loopstandby
The FixerWOLF standby
Findings / Retest0 open
Repro Packs0 packs
Pressure Paths0 paths
Next Movesno runbook
Learning Capsuleproposal gate
Mission Contractplain text -> route
Agent Lanesbounded delegation
Codexcode, repo audit, patch planreceipts
Claudelong-context review, report prosesummary
Hermeslocal tools, shell, rangesartifacts
Browservisual checks and app smoke testsscreens
Evidence Gatesbefore claims harden
Scopeowned target, allowed actionsS3R4PH1M
Proofartifact, log, screenshot, diffledger
Judgefalse-positive and impact reviewJUDG3
Receiptactive tools require approvalscopeguard
Rangeoptional replay when neededCRUCIBLE
Blitz Engine
STRIKES: 0HITS: 0CRITS: 0HIT%: 0EVIDENCE: 0
👥
0
🎯
0
🔑
NO
🛠️
OPT
Quick DemoAuthorized targets
Auto-deploys operators, adds a test target, and launches. No API key needed if a local agent is connected.
🌐
testphp.vulnweb.com
Acunetix test site — SQLi, XSS, CSRF
+
🏦
demo.testfire.net
HCL AppScan demo — auth bypass, injection
+
📡
scanme.nmap.org
Nmap official test host — port scanning
+
Target0
No targets
Operators
0
Config
OPSEC Level, Cognitive Mode, and the toggles above tune the client-side reasoning pipeline — backend/keyless missions currently run operators with their configured prompts + defaults.
QualityACTIVE
0
OK
0
REV
0
REJ
0
ESC
0
Targets
0
Creds
0
Access
0/5
Phases
Findings & Loot(0)
ALLCRITHIGHMEDLOWCREDS|
Severity
Type
Finding
Target
Phase
No findings yet. Engage a mission to discover vulnerabilities.
Enter LaunchEsc AbortD Deploy AllK Clear/ Search
Live Scan
Agent Progress
Operators
0 active
Task Queue
0 tasks
Reasoning And Tool Stream
0 events
ScopeGuard
Scope Receipts
Approval Requests
0 receipts
Active Formation
Custom Deployment
0
Units Active
0
Tools Ready
🚨
0
Critical
⚠️
0
High
ℹ️
0
Medium
🔑
0
Credentials
🔓 Findings
No findings yet. Start a mission to collect evidence.
◉
ARSENAL CONTROL
Select, arm, and coordinate operator tools
85+CATALOG
0ACTIVE
14DOMAINS
Sensor PostureStandby
Arm a few tools to shape the hunt profile.
Visible SetAll tools
Scanning catalog...
Coverage0 lanes
Mix domains for cross-boundary hunts.
HandoffNo operator
Pick an operator before assignment.
◉ ACTIVE LOADOUT
Scopereceipt gate
Breadthneeds domains
Proofneeds validator
Chainneeds composer
No tools assigned - click tools below to add to the loadout
💻 Terminal
T3MP3ST Shell
T3MP3ST Framework v1.0.0
Type 'help' for commands
📈 OBSIDIVM
Quick LaunchPick a preset or build your own mix below
Industry BenchmarksReal CTF challenges from Cybench & NYU CTF Bench
Fetches real challenge descriptions, source code & flags from GitHub. Scores are comparable to published results. Cached for 24h.
Adversarial Tactics, Techniques & Common Knowledge
TA0001 Initial Access --
TA0002 Execution --
TA0003 Persistence --
TA0004 Privilege Escalation --
TA0005 Defense Evasion --
TA0006 Credential Access --
TA0007 Discovery --
TA0008 Lateral Movement --
TA0010 Exfiltration --
TA0011 Command & Control --
🐛 CWE Top 25
--
Most Dangerous Software Weaknesses (2024)
CWE-787 Out-of-Bounds Write --
CWE-79 Cross-Site Scripting --
CWE-89 SQL Injection --
CWE-416 Use After Free --
CWE-78 OS Command Injection --
CWE-20 Input Validation --
CWE-125 Out-of-Bounds Read --
CWE-22 Path Traversal --
CWE-352 Cross-Site Request Forgery --
CWE-434 Unrestricted File Upload --
📊 Overall Results
🏆
--
Overall Score
✅
--
Tests Passed
⏱️
--
Avg Time (ms)
🎯
--
Accuracy
Web
Binary
Crypto
Reverse
Forensics
Auto Ops
OWASP
MITRE
CWE
🧠 AI Performance Analysis
💪 Strengths
⚠️ Weaknesses
🔧 Recommended Improvements
⚙️ Suggested Config Changes
Optimized Configuration
✨ Applied Configuration Changes
🔄
Analyzing performance data...
🤖
LLM is optimizing your configuration...
Analyzing recommendations and generating optimal settings
⚠ Local simulation / illustrative. The container controls and challenge "solves" here do not touch a live Docker daemon or a real flag server — flags are format-checked, not verified against a live target. For measured, live-exploit-verified results use OBSIDIVM and the benchmarks.
🚀 Quick Setup
Get started with execution-based CTF benchmarks in minutes. Requires Docker installed locally.
1️⃣Check DockerUnknown
Verify Docker daemon is running and accessible.
2️⃣Build ImagesNot Built
Build Docker images for CTF challenges.
3️⃣Launch RangeOffline
Start all challenge containers.
📋 Manual Commands (run in terminal)
cd ctf && docker-compose build && docker-compose up -d
🎯 Challenge Browser
📊 Range Status
Containers Running0
Challenges Solved0 / 8
Total Points0 / 1450
Agent Success Rate--%
🤖 Agent Execution
Initializing agent...0 / 8
📈 Results & History
Challenge
Category
Agent
Time
Status
Flag
🏁
No execution results yet. Run the benchmark to see agent performance.
★ THE ADMIRAL — Autonomous Op Orchestrator
STANDING BY
Give the Admiral a high-level directive and it will autonomously plan the entire operation — identifying targets, allocating operators, setting OPSEC levels, and executing the full kill chain. No manual configuration needed.
Operation Plan
Admiral's Intent
UNROUTED
Mission Gate
WAIT
Targets
Force Allocation
Objectives
Rules of Engagement
Hunt Lanes
Specialist Work Orders
Evidence Contract
Critic & Tool Route
Admiral's Strategic Rationale
Situation Reports
Awaiting first situation report...
Strategic Assessment
Assessment will be generated after operation completes or on demand.
Loading self-improvement loops…
🎛️ Universal API Config
Configure your LLM providers in one place — provider, key, base URL, model, and context cap — shared with the individual provider sections below (fully backward-compatible).
In-browser mission routing currently uses OpenRouter or Venice (or a connected local agent). Keys for other providers, custom base URLs, and the context cap are stored and used for model discovery and server-side / agent flows.
Stored, but direct-LLM currently routes via OpenRouter or a connected agent — not wired to this key yet.
OpenAI API Key
Stored, but direct-LLM currently routes via OpenRouter or a connected agent — not wired to this key yet.
🖥️ Local Model (llama.cpp / Ollama / OpenAI-compatible)
Point the War Room at a model running on your own host — no cloud, no OpenRouter.
Requires npm run server. When enabled, BOTH server-dispatched missions/General AND the
browser "AI" features route through your local model.
Route all probe/attack fetch() traffic through a SOCKS5 proxy so tests
don't leave from your own IP (Tor, an SSH -D tunnel, a VPS, etc.). Loopback (the local model & this
server) is always bypassed. Requires npm run server; also settable via TEMPEST_PROXY_URL.
Tor: socks5://127.0.0.1:9050 · SSH tunnel: ssh -D 1080 host → socks5://127.0.0.1:1080. Use socks5h:// to resolve DNS at the proxy.
🔌 Local Agents
Enlist agents you've already authed on this machine — no API keys needed. T3MP3ST detects each CLI (Claude Code / Codex / Hermes) and drives it as your LLM backbone using its own login. Connect one, then hit ☆ Use to pin it as the ★ active backbone for missions and analysis — a pinned local model outranks any stored API key.
🔌 Connect Local Agents
— scanning —
Detecting local agents…
🟠
Claude Code model
Which local Claude model to run (passed to claude --model)
🤖 Model Selection
Select your preferred AI model for agent operations
Loading models...
🔄 Fallback ModelAuto-retries with this model if the primary fails (refusal, timeout, error)
🖥️ API Server
URL of your running T3MP3ST API server (start with npm run server)
Default: http://localhost:3333 — change if your server runs on a different host/port
⚠️ Data
Config Library
Saved Configurations
0 configs saved •
Default: None
📊 Current Configuration
Not benchmarked
Run a benchmark to see current config performance
💾 Saved Configurations
💾
No saved configurations yet
Run benchmarks and save your best performing configs