MB_CORE_LOG
Agent Memory & RAGStatus: Prototype

Adaptive Runbook Intelligence Platform

A support-automation proof of concept that learns reusable runbooks from execution history.

Tech Stack: Python · ChromaDB · SQLite · Streamlit · MCP · Embeddings

Problem

Repeated support requests can trigger the same retrieval and reasoning even after a successful resolution is already known.

What I built

Built knowledge retrieval, case memory, a runbook library, policy-based routing, feedback tracking, and a fast path for known-good runbooks without LLM calls.

Key decisions

Reuse explicit runbooks: Stored steps can be inspected and routed directly when their history supports reuse. Separate semantic and structured memory: Similarity search finds candidates while counters support lifecycle decisions.

Results

Includes a three-phase benchmark over 20 synthetic support tickets, comparing stateless reasoning, runbook creation, and runbook reuse.

Limitations

The demonstration uses synthetic tickets and simulated MCP actions. It does not establish production incident-resolution performance.

Evidence

The repository includes workflow code, runbook storage, policy logic, a benchmark runner, and documented metrics.

1. Problem Statement

Repeated support requests can trigger the same retrieval and reasoning even after a successful resolution is already known.

2. Real-World Motivation

Reuse validated resolution steps and retain feedback about when a runbook succeeds or fails.

3. System Architecture

Two execution paths share local memory: reuse an eligible runbook or retrieve context for a new LLM proposal.

The repository is a local proof of concept. MCP actions are simulated Python functions; enterprise systems are not connected. Dashed arrows show persisted updates, with new candidates and cases created on the exploratory path.

4. Pipeline Data Flow

A request bypasses model reasoning only after the runbook and reuse policy checks pass.

This sequence follows RUNBOOK_AWARE mode. The source requires a known-good match at similarity 0.75 or higher and a reuse policy score of at least 0.55; a rejected reuse gate falls back to exploratory reasoning.

5. Failure Modes & Mitigations

ScenarioImpactMitigation Strategy
No suitable runbookA saved resolution cannot be reused confidently.Route through exploratory retrieval and reasoning.
Runbook repeatedly fails or reopensA previously useful procedure may no longer be reliable.Update lifecycle counters and exclude known-bad runbooks from matching.

6. Design Tradeoffs

DecisionAlternativeRationale
Reuse explicit runbooksReason from scratch for every requestStored steps can be inspected and routed directly when their history supports reuse.
Separate semantic and structured memoryStore all runbook state in one formatSimilarity search finds candidates while counters support lifecycle decisions.

7. Validation

The benchmark reuses a fixed synthetic ticket set across three phases and includes a repeated-query determinism check.

8. Setup & Delivery

The README documents Python setup, provider configuration, benchmark commands, and Streamlit launch instructions.

Results & Evaluation

Evaluation Summary
The known-good fast path bypasses model reasoning and tracks latency, tokens, and LLM calls separately.
Evaluation Scope
  • Three phases compare stateless, learning, and reuse behavior.
  • Repeated-query checks compare determinism hashes.

Scaling Strategy

The proof of concept uses ChromaDB indexes and SQLite statistics. Distributed operation is not demonstrated.

Security Model

A policy gate evaluates execution eligibility. Benchmark MCP actions are simulated rather than live enterprise changes.

Observability

  • Execution metrics
    Tokens, model calls, latency, and escalations.
  • Runbook lifecycle
    Successes, failures, reopenings, and reuse status.

Future Roadmap

Further validation would use held-out cases and approved integrations with real operational systems.