A Controlled Benchmark of Edge AI Accelerators for Traffic Object Detection

avoin
Julkaisu on tekijänoikeussäännösten alainen. Teosta voi lukea ja tulostaa henkilökohtaista käyttöä varten. Käyttö kaupallisiin tarkoituksiin on kielletty.
Lataukset1

Verkkojulkaisu

DOI

Tiivistelmä

Stringent power and cost constraints have driven the adoption of purpose-built edge accelerators over conventional processors for real-time visual perception, yet archi- tectural heterogeneity prevents direct comparison of published specifications for these accelerators. To resolve this, this thesis presents a controlled benchmark of four edge-inference paradigms: a fixed-function TPU (Coral Dev Board Mini), an SoC NPU (Khadas VIM3), an embedded GPU (NVIDIA Jetson Nano), and a reconfigurable FPGA (AMD Kria KV260). Under a unified evaluation protocol, an INT8 EfficientDet-Lite2 detector, trained on a seven-class BDD100K subset, was deployed across the platforms, with a YOLOv4-416 substitute on the GPU due to toolchain constraints. Each platform was measured on latency, accuracy, power, energy, and cost. Results demonstrate that no single paradigm dominates, and rated throughput is an unreliable performance predictor. Despite its 5.0 TOPS rating, the NPU was the slowest and least energy-efficient platform, yielding 0.71 frames per second and consuming 3.84 J per inference. Conversely, despite its lower 4.0 TOPS rating, the TPU ran 4.2 times faster per inference, delivered 6.8 times higher throughput per watt at 1.77 frames per second per watt, and at 91.99 C was the most cost-effective. The two fixed-function accelerators agreed within 0.001 mAP, confirming consistency. However, FPGA deployment collapsed accuracy from 0.2400 to 0.0233. A calibration ablation and bias-bypass experiment isolated this degradation to the per-tensor power-of-two quantization within the classification head rather than the hardware itself.

item.page.okmtext