Run the same prompt across Claude, GPT, Gemini, and your local models at once. See every answer side by side, get one score for how much they agree, and catch the failures before your users do — privately, on your own hardware.
Most teams never find out until production — when a customer does.
Amagra shows you all four before deployment — running your prompt everywhere at once and scoring exactly where the models disagree.
You already know this loop from code: install, paste, run, see the errors, fix, run again. Amagra is that exact loop for the prompts you write — same muscle memory, no cloud.
| The loop | VS Code + Python | Amagra |
|---|---|---|
| Get it | Installer · ~300 MB · ~5 min | docker pull d4shm1r/amagra |
| Runtime | Install the Python extension first | ✓ Local model included — zero key, offline |
| New file | Open test.py, paste code | Open the debugger, paste your prompt |
| See problems | Hit Run → 10 errors print | Before you run — health score + role / task / format checks |
| Fix | Edit all 10 by hand | ✓ One-click auto-repair, checks turn green live |
| Run | ▶ → one output | Run across models — Claude / GPT / local, side by side |
| Read result | Output is output | Outputs + latency + a divergence verdict |
VS Code tells you your code is broken. Amagra tells you your prompt is under-specified — and runs it across every model so you see exactly where they disagree.
Paste once, run everywhere, read one answer instead of three.
An open architecture with a documented API — see exactly how every decision was made.
Debugging prompts is the way in. Once you're inside, Amagra keeps working for you — routing, remembering, and recording every decision it makes.
One download, double-click, done — the backend and UI ship together and run entirely on your machine.
Unsigned for now, so first launch shows a Gatekeeper / SmartScreen prompt — open anyway. Prefer to run with Docker or build from source →
The self-hosted version is free forever. Managed plans add hosting and a frontier-model backend — no GPU required.
Privacy, supported models, pricing, and exactly how the divergence score is computed — all answered in full on the FAQ page.
Read the FAQ →The self-hosted version is free forever and runs on hardware you already own.