Adaptive Compute Efficient Learning via Conceptual-Criticality
Research prototypes for entropy-based difficulty prediction and confidence-based early-exit inference.
Tech Stack: Python · PyTorch · Transformers · Jupyter · Early Exit
Problem
A fixed inference depth spends the same computation on easy and difficult inputs.
What I built
Co-authored the AAAI 2026 Student Abstract and contributed to separate notebook prototypes for criticality estimation and early-exit inference.
Key decisions
Learn an entropy-based difficulty predictor: A pooled-embedding MLP learns difficulty buckets derived from LSTM prediction entropy in a separate experiment. Use confidence-based early exits: Threshold selection makes the accuracy-versus-compute tradeoff explicit.
Results
Resume-reported proof-of-concept result: about 90.7% accuracy and about 65% lower energy use versus a 6-layer baseline. The notebook estimates energy and extrapolates baseline cost.
Limitations
Criticality prediction and early exit are separate notebook experiments. Energy is estimated from sampled GPU power and elapsed time; six-layer baseline cost is extrapolated.
Evidence
The research repository contains the paper, notebooks, and experiment code. Results here follow the resume summary.
1. Problem Statement
A fixed inference depth spends the same computation on easy and difficult inputs.
2. Real-World Motivation
Study entropy-based input difficulty and confidence-based early exit in separate prototypes while tracking prediction quality.
3. System Architecture
The criticality predictor is not connected to the early-exit loop in these notebooks. Training creates model weights; the early-exit experiment then sweeps inference thresholds against a full-depth baseline on AG News.
4. Pipeline Data Flow
This follows predict_until_exit with batch size one. Energy is estimated from a sampled GPU power reading multiplied by elapsed time. The notebook extrapolates baseline latency and energy; it does not directly measure those baseline costs.
5. Failure Modes & Mitigations
| Scenario | Impact | Mitigation Strategy |
|---|---|---|
| An input exits too early | Compute savings may reduce prediction quality. | Evaluate accuracy alongside exit thresholds and compute usage. |
| Results vary across workloads | Savings may not transfer to another model or dataset. | Report the evaluated baseline and experimental scope. |
6. Design Tradeoffs
| Decision | Alternative | Rationale |
|---|---|---|
| Learn an entropy-based difficulty predictor | Use manually assigned difficulty labels | A pooled-embedding MLP learns difficulty buckets derived from LSTM prediction entropy in a separate experiment. |
| Use confidence-based early exits | Always use the final layer | Threshold selection makes the accuracy-versus-compute tradeoff explicit. |
7. Validation
Notebooks demonstrate criticality estimation and early-exit comparisons alongside the research paper and experiment code.
8. Setup & Delivery
Python dependency setup and notebook instructions are available for reproducing the research workflow.
Results & Evaluation
- Baseline: extrapolated cost for fixed six-layer computation.
- Evaluation: accuracy, exit layer, elapsed time, and GPU power-based energy estimates.
Scaling Strategy
Separate notebook experiments explore difficulty prediction and early-exit inference. An integrated allocator and larger-workload validation remain further research.
Security Model
This is experimental research code rather than an exposed inference service; no service security model is claimed.
Observability
- AccuracyPredictive quality in the evaluated experiment.
- Compute useEnergy, layers used, and inference-cost comparisons.
Future Roadmap
Further research would test more workloads and document how savings vary with model size and exit thresholds.