This monograph addresses the critical efficiency gap between the theoretical accuracy of multimodal deep neural networks and their deployment cost on resource-constrained edge platforms. It identifies a common root cause—multi-layered "System Rigidity"—that leads to the crisis of "Dark Intelligence," where edge devices cannot sustain their theoretical AI capabilities due to strict thermal and power limits.
To dismantle this rigidity, the book pioneers the paradigm of Energy-Proportional Intelligence (EPI), ensuring that a system's energy consumption scales strictly in proportion to the intrinsic cognitive complexity of the input data. The book presents a comprehensive, full-stack approach: it begins with a novel benchmark suite (MMBench) that reveals physical workload characteristics and bottlenecks. It then introduces model-layer adaptive inference techniques (MMExit, MMBypass) that dynamically reduce computation, scheduling-layer runtime optimizations (CPM, A²) that manage power and heterogeneous accelerators under strict constraints, and architecture-layer modality gating designs (SMG, AMG) that proactively cut sensing and preprocessing costs.
Readers will gain a cohesive cross-layer methodology and practical techniques to make multimodal systems accurate, fast, and energy-efficient on edge devices. Prerequisites include a graduate-level familiarity with computer architecture, operating systems, and machine learning.
This monograph addresses the critical efficiency gap between the theoretical accuracy of multimodal deep neural networks and their deployment cost on resource-constrained edge platforms. It identifies a common root cause—multi-layered "System Rigidity"—that leads to the crisis of "Dark Intelligence," where edge devices cannot sustain their theoretical AI capabilities due to strict thermal and power limits.
To dismantle this rigidity, the book pioneers the paradigm of Energy-Proportional Intelligence (EPI), ensuring that a system's energy consumption scales strictly in proportion to the intrinsic cognitive complexity of the input data. The book presents a comprehensive, full-stack approach: it begins with a novel benchmark suite (MMBench) that reveals physical workload characteristics and bottlenecks. It then introduces model-layer adaptive inference techniques (MMExit, MMBypass) that dynamically reduce computation, scheduling-layer runtime optimizations (CPM, A²) that manage power and heterogeneous accelerators under strict constraints, and architecture-layer modality gating designs (SMG, AMG) that proactively cut sensing and preprocessing costs.
Readers will gain a cohesive cross-layer methodology and practical techniques to make multimodal systems accurate, fast, and energy-efficient on edge devices. Prerequisites include a graduate-level familiarity with computer architecture, operating systems, and machine learning.
Xiaofeng Hou
Dark Intelligence Multimodal Computing Deep Neural Networks Dynamic Inference Early Exiting DVFS Heterogeneous accelerators modality Gating Edge AI Power Management Hardware-Software Co-Design Energy-Proportional Intelligence