A structured study reference covering 9 core topics for thesis preparation on perception and decision-making for autonomous vehicles at road signals. The document is organized for students or researchers building a thesis pipeline, with each topic following a consistent format: plain explanation, importance, fundamental techniques, alternatives, and connections to other topics.
This document is a structured study reference for thesis preparation, covering 9 core topics on perception and decision-making for autonomous vehicles at road signals. It is intended for students or researchers building a pipeline for autonomous driving, with each topic following a consistent format: plain explanation, importance, fundamental techniques, alternatives, and connections to other topics.
The document is divided into two layers. Layer A covers perception, or seeing the road, and includes topics 1 through 6. These cover OpenCV as a toolkit for image processing, YOLO as a real-time object detector, DETR as a transformer-based detection approach, Grounding DINO for text-prompted detection, DINO for self-supervised pretraining, and obstacle detection as the applied task.
Layer B covers decision-making, or acting on the road, and includes topics 7 through 9. These cover distance calculation from 2D boxes to real-world numbers, speed calculation, and path planning and motion planning. The document explains how these layers connect, with cameras and detectors telling the car what is around it and motion planning using that to decide brake, turn, or continue.
The document provides detailed comparisons between detection models, including YOLO variants, RT-DETR, and RF-DETR, with specific performance metrics such as F1-scores and mAP. It also discusses practical considerations like the trade-off between speed and accuracy, the use of zero-shot detection for flexibility, and the importance of robust performance under difficult environmental conditions.
For distance calculation, the document covers three families of techniques: sensor-based methods using LiDAR or radar, monocular vision using a single camera, and binocular or stereo vision using two cameras. It notes typical error rates and mentions newer approaches like deep-learning monocular depth estimation and fiducial-marker methods that reduce estimation noise.