MB_CORE_LOG
AI ResearchStatus: Prototype

Adaptive Compute Efficient Learning via Conceptual-Criticality

Research prototypes for entropy-based difficulty prediction and confidence-based early-exit inference.

Tech Stack: Python · PyTorch · Transformers · Jupyter · Early Exit

Problem

A fixed inference depth spends the same computation on easy and difficult inputs.

What I built

Co-authored the AAAI 2026 Student Abstract and contributed to separate notebook prototypes for criticality estimation and early-exit inference.

Key decisions

Learn an entropy-based difficulty predictor: A pooled-embedding MLP learns difficulty buckets derived from LSTM prediction entropy in a separate experiment. Use confidence-based early exits: Threshold selection makes the accuracy-versus-compute tradeoff explicit.

Results

Resume-reported proof-of-concept result: about 90.7% accuracy and about 65% lower energy use versus a 6-layer baseline. The notebook estimates energy and extrapolates baseline cost.

Limitations

Criticality prediction and early exit are separate notebook experiments. Energy is estimated from sampled GPU power and elapsed time; six-layer baseline cost is extrapolated.

Evidence

The research repository contains the paper, notebooks, and experiment code. Results here follow the resume summary.

1. Problem Statement

A fixed inference depth spends the same computation on easy and difficult inputs.

2. Real-World Motivation

Study entropy-based input difficulty and confidence-based early exit in separate prototypes while tracking prediction quality.

3. System Architecture

Separate notebook prototypes study entropy-based difficulty prediction and confidence-based early-exit classification.

The criticality predictor is not connected to the early-exit loop in these notebooks. Training creates model weights; the early-exit experiment then sweeps inference thresholds against a full-depth baseline on AG News.

4. Pipeline Data Flow

Classifier heads at layers 2, 4, and 6 decide whether a sample can return a prediction before full-depth computation.

This follows predict_until_exit with batch size one. Energy is estimated from a sampled GPU power reading multiplied by elapsed time. The notebook extrapolates baseline latency and energy; it does not directly measure those baseline costs.

5. Failure Modes & Mitigations

ScenarioImpactMitigation Strategy
An input exits too earlyCompute savings may reduce prediction quality.Evaluate accuracy alongside exit thresholds and compute usage.
Results vary across workloadsSavings may not transfer to another model or dataset.Report the evaluated baseline and experimental scope.

6. Design Tradeoffs

DecisionAlternativeRationale
Learn an entropy-based difficulty predictorUse manually assigned difficulty labelsA pooled-embedding MLP learns difficulty buckets derived from LSTM prediction entropy in a separate experiment.
Use confidence-based early exitsAlways use the final layerThreshold selection makes the accuracy-versus-compute tradeoff explicit.

7. Validation

Notebooks demonstrate criticality estimation and early-exit comparisons alongside the research paper and experiment code.

8. Setup & Delivery

Python dependency setup and notebook instructions are available for reproducing the research workflow.

Results & Evaluation

Evaluation Summary
Reported proof-of-concept figures: about 90.7% accuracy and about 65% estimated energy reduction versus the six-layer baseline.
Evaluation Scope
  • Baseline: extrapolated cost for fixed six-layer computation.
  • Evaluation: accuracy, exit layer, elapsed time, and GPU power-based energy estimates.

Scaling Strategy

Separate notebook experiments explore difficulty prediction and early-exit inference. An integrated allocator and larger-workload validation remain further research.

Security Model

This is experimental research code rather than an exposed inference service; no service security model is claimed.

Observability

  • Accuracy
    Predictive quality in the evaluated experiment.
  • Compute use
    Energy, layers used, and inference-cost comparisons.

Future Roadmap

Further research would test more workloads and document how savings vary with model size and exit thresholds.