Capability Runner
Discover browser workflows once, then replay them with new inputs without calling a model.
Tech Stack: Python · Playwright · React · TypeScript · LLM Tool Calling
Problem
Repeating model-driven browser reasoning for familiar tasks adds cost and makes outcomes harder to reproduce.
What I built
Built LLM-assisted workflow discovery, saved capability artifacts, deterministic replay, action validation, page-evidence checks, expected-error handling, and human takeover.
Key decisions
Discover once, replay deterministically: A saved procedure makes repeated execution inspectable and avoids repeated model decisions. Validate state before continuing: Page evidence must support the requested result.
Results
Recorded local banking demos show saved workflows replaying with different inputs and zero replay model calls.
Limitations
Demonstrated on local banking apps with synthetic data. General website support, production authentication, and cross-tenant reuse are outside the demonstrated scope.
Evidence
Repository, design report, screenshot guide, and recorded discovery/replay evidence are available on GitHub.
1. Problem Statement
Repeating model-driven browser reasoning for familiar tasks adds cost and makes outcomes harder to reproduce.
2. Real-World Motivation
Separate learning a workflow from executing it, while checking that each saved action still matches the current page.
3. System Architecture
Source-reviewed implementation view at commit 1b7eea0. Arrows show the principal runtime dependencies; both engines inspect fresh browser observations. Demonstrated on local banking apps with synthetic data. Sessions are process-local, and physical human-handoff acceptance remains outstanding.
4. Pipeline Data Flow
Replay path condensed from the implementation at commit 1b7eea0. The loop shows authorized actions; blocked or uncertain actions stop execution or enter the explicit intervention path. A successful click alone does not establish task completion.
5. Failure Modes & Mitigations
| Scenario | Impact | Mitigation Strategy |
|---|---|---|
| Stale or ambiguous page target | A saved action may point at the wrong control. | Validate fresh evidence and fail when target resolution is uncertain. |
| Action needs human intervention | Replay cannot safely continue automatically. | Pause for takeover and validate fresh state before resuming. |
6. Design Tradeoffs
| Decision | Alternative | Rationale |
|---|---|---|
| Discover once, replay deterministically | Ask a model to plan every run | A saved procedure makes repeated execution inspectable and avoids repeated model decisions. |
| Validate state before continuing | Treat completed clicks as success | Page evidence must support the requested result. |
7. Validation
Tests and evaluation commands cover replay, expected business errors, and intervention behavior. The design report links recorded evidence.
8. Setup & Delivery
The repository documents Python and frontend setup, local launch commands, and evaluation instructions.
Results & Evaluation
- Recorded replay checks expected outputs with new inputs.
- Evidence covers local demo cases, not arbitrary websites.
Scaling Strategy
One Python process coordinates the browser. Active sessions are local and do not survive a restart.
Security Model
The action gateway checks allowed actions, ownership, and observations. The local UI is not a production authentication system.
Observability
- Execution evidenceSaved steps, outcomes, and failure context.
- Model callsDiscovery and replay usage are tracked separately.
Future Roadmap
Further validation would cover physical handoff acceptance and reuse of one unchanged capability across application profiles.