# Self-Healing Agent

An AI agent that woke up before you did. It monitored production, diagnosed errors, and opened PRs with fixes—all autonomously.

Source: https://miloscvetkovic.dev/work/self-healing-agent

- Tagline: autonomous bug fixing
- Category: AI AGENT
- Status: RETIRED
- Metric: 73% errors resolved autonomously
- Basis: Production errors the agent diagnosed and fixed in a pull request the team merged, out of all the production errors it monitored.
- Tags: Claude Agent SDK, Bun, Elysia, Azure
- Published: 2026-09-09
- Updated: 2026-10-01

## The Challenge

Production breaks at 3am, and nobody wants that call. Errors don't wait for business hours, every minute of downtime costs money and trust, and the on-call rotation turns into the job developers dread. Before anyone can fix anything, somebody has to wake up, read the production logs, find their way around the codebase and work out what actually broke. That diagnosis is exactly the kind of grunt work a machine could do overnight, if it could be trusted with it. Handing it to an agent brings risks of its own, though: an agent can be confidently wrong, it can run up a bill, and it can keep trying long after a sensible engineer would have stopped. The question: can we fix bugs faster than humans can even wake up?

## My Approach

I built an autonomous agent that monitored production errors and proposed fixes through pull requests. It ran on Bun and Elysia, kept its data in Azure Table Storage and used Azure Log Analytics for monitoring. When something broke, an error analysis pipeline built on the Claude Agent SDK read the failure against the codebase, diagnosed the issue and drafted a fix, which the agent opened as a pull request through the GitHub API. It monitored the CI pipeline and retried on failure, capped at three attempts, so a stubborn failure could not loop forever. Safety constraints kept it on a short leash: daily limits, budget caps and confidence thresholds bounded what it could attempt, and it had a kill switch and emergency override controls. It tracked learning metrics to improve its calibration. Humans reviewed and merged its pull requests through approval gates; the agent did the grunt work. The agent has since been retired.

## How It Works

1. The agent ran on Bun and Elysia and monitored production errors through Azure Log Analytics.
2. When something broke, an error analysis pipeline built on the Claude Agent SDK read the error against the codebase and diagnosed the issue.
3. It drafted a fix and opened a pull request through the GitHub API, while safety constraints (daily limits, budget caps and confidence thresholds) bounded what it could attempt.
4. It also monitored the CI pipeline and retried on failure, up to three attempts.
5. Humans reviewed the pull request through approval gates and merged the fix, with a kill switch and emergency override controls on hand.
6. It tracked learning metrics to improve its calibration.

## Key Contributions

- Designed autonomous error analysis pipeline using Claude AI
- Implemented automatic PR creation with contextual fixes
- Built CI pipeline monitoring with retry logic (max 3 attempts)
- Added safety constraints: daily limits, budget caps, confidence thresholds
- Created kill switch and emergency override controls
- Implemented learning metrics for calibration improvement

## Impact

- Production incidents got diagnosed before anyone woke up
- Developers stopped dreading on-call rotations
- Fix patterns got learned and reused automatically
- Humans stayed in control—agent proposed, team approved
- The system literally improved itself over time

## Tech Stack

| Category | Items |
| --- | --- |
| Runtime | Bun |
| Framework | Elysia |
| AI | Claude Agent SDK, Anthropic API |
| Monitoring | Azure Log Analytics |
| Storage | Azure Table Storage |
| VCS | GitHub API |
