MB_CORE_LOG

Selected Engineering Projects

Backend systems, AI-agent workflows, and research experiments, with implementation details and project evidence.

Browser AutomationStatus: Prototype

Capability Runner

Python · Playwright · React · TypeScript · LLM Tool Calling

Discover browser workflows once, then replay them with new inputs without calling a model.

Problem: Repeating model-driven browser reasoning for familiar tasks adds cost and makes outcomes harder to reproduce.
What I built: Built LLM-assisted workflow discovery, saved capability artifacts, deterministic replay, action validation, page-evidence checks, expected-error handling, and human takeover.
Key decisions: Discover once, replay deterministically; Validate state before continuing.
Results: Recorded local banking demos show saved workflows replaying with different inputs and zero replay model calls.
Limitations: Demonstrated on local banking apps with synthetic data. General website support, production authentication, and cross-tenant reuse are outside the demonstrated scope.
Evidence: Repository, design report, screenshot guide, and recorded discovery/replay evidence are available on GitHub.
Distributed SystemsStatus: Prototype

SyncStream

Java · PostgreSQL · Debezium · Kafka · Redis · Elasticsearch · Docker

Distributed data synchronization using PostgreSQL change events, Kafka, Redis, and Elasticsearch.

Problem: Synchronizing caches and search indexes through application dual writes creates failure windows and duplicated integration logic.
What I built: Built a CDC pipeline with Debezium, Kafka consumers for Redis and Elasticsearch, shared retries, dead-letter queues, and consumer-management APIs.
Key decisions: CDC through Debezium and Kafka; Separate cache and search consumers.
Results: Implemented product-event propagation from PostgreSQL to cache and search projections, with repository verification scripts.
Limitations: A local development platform; analytics consumption is not implemented and role-header authorization needs production hardening.
Evidence: Source, Docker Compose setup, architecture documentation, and workflow verification scripts are available in the repository.
Multimodal AI AgentsStatus: Prototype

Ticketless IT/HR Voice Support

Speech Recognition · WebSockets · VLM/OCR · RAG · LangGraph

A voice and screen support prototype combining enterprise retrieval, guided actions, and human approval.

Problem: Support requests often need both a spoken description and the context visible on an employee's screen.
What I built: Built speech recognition, WebSocket voice interactions, VLM/OCR screen understanding, enterprise RAG, and LangGraph orchestration with human approval and structured verification before closure.
Key decisions: Use voice and screen context; Keep human approval.
Results: Implemented a combined voice-and-screen support workflow. No quantitative performance result is claimed for this prototype.
Limitations: An IT/HR support prototype; production deployment and a measured reliability benchmark are not established by this project summary.
Evidence: Implementation summary follows my resume. The supplied project repository is linked above; public access is currently unavailable.
Agent Memory & RAGStatus: Prototype

Adaptive Runbook Intelligence Platform

Python · ChromaDB · SQLite · Streamlit · MCP · Embeddings

A support-automation proof of concept that learns reusable runbooks from execution history.

Problem: Repeated support requests can trigger the same retrieval and reasoning even after a successful resolution is already known.
What I built: Built knowledge retrieval, case memory, a runbook library, policy-based routing, feedback tracking, and a fast path for known-good runbooks without LLM calls.
Key decisions: Reuse explicit runbooks; Separate semantic and structured memory.
Results: Includes a three-phase benchmark over 20 synthetic support tickets, comparing stateless reasoning, runbook creation, and runbook reuse.
Limitations: The demonstration uses synthetic tickets and simulated MCP actions. It does not establish production incident-resolution performance.
Evidence: The repository includes workflow code, runbook storage, policy logic, a benchmark runner, and documented metrics.
AI ResearchStatus: Prototype

Adaptive Compute Efficient Learning via Conceptual-Criticality

Python · PyTorch · Transformers · Jupyter · Early Exit

Research prototypes for entropy-based difficulty prediction and confidence-based early-exit inference.

Problem: A fixed inference depth spends the same computation on easy and difficult inputs.
What I built: Co-authored the AAAI 2026 Student Abstract and contributed to separate notebook prototypes for criticality estimation and early-exit inference.
Key decisions: Learn an entropy-based difficulty predictor; Use confidence-based early exits.
Results: Resume-reported proof-of-concept result: about 90.7% accuracy and about 65% lower energy use versus a 6-layer baseline. The notebook estimates energy and extrapolates baseline cost.
Limitations: Criticality prediction and early exit are separate notebook experiments. Energy is estimated from sampled GPU power and elapsed time; six-layer baseline cost is extrapolated.
Evidence: The research repository contains the paper, notebooks, and experiment code. Results here follow the resume summary.
Retrieval & EvaluationStatus: Prototype

Dynamic vs. Fixed K Reranking in RAG

Python · Sentence Transformers · FAISS · Cross-Encoder · Jupyter

A retrieval experiment comparing fixed context size with relevance-based dynamic document selection.

Problem: A fixed document count can add irrelevant context for simple queries or omit context for broader queries.
What I built: Built a notebook comparison using embeddings, FAISS retrieval, cross-encoder reranking, and fixed versus score-based dynamic context selection.
Key decisions: Select by relative reranker score; Evaluate retrieval independently.
Results: In the documented synthetic experiment, mean context word count fell from 125.17 to 13.00 while answer presence in context remained 1.00 for both methods.
Limitations: Uses 100 synthetic documents and 30 queries. The notebook uses word count as a token-cost proxy; answer presence measures retrieved context, not generated-answer correctness.
Evidence: The repository contains the comparison notebook, experiment design, evaluation definitions, and aggregate results.