Skip to main content

Multimodal sensing and perception for autonomous vehicles

In a nutshell, the objective of the project is to push the boundaries in the field of training DNNs for perception from multimodal imaging sensors, based on datasets captured in challenging environmental conditions (bad weather, night time) and applied to the detection of vulnerable road users (VRUs), including pedestrians, cyclists, and (for the first time) micro-mobility users. With this purpose, we will integrate and calibrate the sensors on a vehicle for data acquisition, developing an ad-hoc interface for data visualization, flow, and storage. The vehicle will capture and store the data which, once annotated in the different categories, will constitute the datasets for the project that will be made publicly available, and which will be used to explore different data fusion and training strategies. Relevant secondary objectives, such as the analysis of the limits of each imaging mode, an annotation tool for multimodal images, a procedure for hardware integration In a multimodal vehicle, the creation of the first dataset specifically involving micro-mobility users, the analysis of early and late fusion algorithms using different imaging modes, and the exploration of state of the art multimodal DNNs, are all of them direct contributions to the accuracy and reliability of AVs which will create tools for advancing the field.

Accordingly, the project has been divided into three work packages covering each of these main groups of tasks (hardware integration of a multimodal data collection unit, dataset generation and publication, and DNN training using multimodal data), plus an initial one for defining in detail the methods of the project and provisioning the sensors. The research team, as required by the multidisciplinary effort tackled,  is diverse as it includes specialists in hardware (photonics, engineers) and software(telecommunication engineers, MSc in computer Vision), both with senior (2PhD) and junior (3PhD students) profiles. The project builds on the expertise in system integration, lidar imaging, and data fusion already created within UPC jointly by a stable collaboration of CD6 (the Optical Engineering and Sensors group led by the IP, Santiago Royo) with GPI (the image processing group led by Josep Ramon Casas), both members of the research team (Equipo de Investigación) of the project, yielding a truly multidisciplinary project where software is dependent on the quality of the data generated by hardware, but hardware is useless without an effective perception which enables decision-making in the AV.

Project TED2021-132338B-I00, funded by MCIU/AEI/10.13039/501100011033 and by the European Union Next Generation EU/PRTR