On-device AI execution lacks unified runtime observability across mobile platforms. Existing tools focus on isolated model profiling or single-engine metrics and give no synchronized view of memory, thermal state, battery draw, and inference latency during real workloads. We present EdgePulse, an open-source runtime observability framework for on-device AI. EdgePulse captures structured InferenceTrace telemetry on Android (Kotlin) and iOS (Swift) via native platform channels: resident memory via Debug.getPss(), OS thermal state via PowerManager.currentThermalStatus, battery current draw via BatteryManager.CURRENT_NOW, CPU utilization from /proc/stat, and per-inference latency. We evaluate EdgePulse through a real-device empirical study on a physical Tecno CH7n (MediaTek Helio G35, 4 GB RAM, Android 12, API 31) across three runtimes: TFLite (MobileNet-V3-Small), ONNX Runtime (ResNet18), and GGUF/llama.cpp (TinyLlama 1.1B Q4). Analyzing 90 baseline traces and 60 sustained-load traces reveals: baseline mean latency varied 12× across runtimes (148 ms to 1,779 ms); sustained MobileNet inference over 15 minutes showed no degradation; sustained TinyLlama decoding produced a +16.5% latency increase (1,829 ms → 2,131 ms) while PowerManager.currentThermalStatus reported nominal throughout. This thermal-API blind spot demonstrates that standard OS thermal APIs can miss active hardware throttling on budget ARM SoCs, making application-level timing essential for edge AI deployments.
No takes yet. Share an insight, caveat, or question.
Muhammad Assad Ullah (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: