Post-Doctoral Research Visit F - M Continual Audiovisual Perception For Human-Robot Interaction H/F - INRIA
- Villé - 67
- CDD
- INRIA
Les missions du poste
A propos d'Inria Inria est l'institut national de recherche dédié aux sciences et technologies du numérique. Il emploie 2600 personnes. Ses 215 équipes-projets agiles, en général communes avec des partenaires académiques, impliquent plus de 3900 scientifiques pour relever les défis du numérique, souvent à l'interface d'autres disciplines. L'institut fait appel à de nombreux talents dans plus d'une quarantaine de métiers différents. 900 personnels d'appui à la recherche et à l'innovation contribuent à faire émerger et grandir des projets scientifiques ou entrepreneuriaux qui impactent le monde. Inria travaille avec de nombreuses entreprises et a accompagné la création de plus de 200 start-up. L'institut s'eorce ainsi de répondre aux enjeux de la transformation numérique de la science, de la société et de l'économie.
Post-Doctoral Research Visit F/M Continual Audiovisual Perception for Human-Robot Interaction
Le descriptif de l'offre ci-dessous est en Anglais
Type de contrat : CDD
Niveau de diplôme exigé : Thèse ou équivalent
Fonction : Post-Doctorant
A propos du centre ou de la direction fonctionnelle
The Centre Inria de l'Université de Grenoble groups together almost 450 people in 26 research teams and 9 research support departments.
Staff is present on three campuses in Grenoble, in close collaboration with other research and higher education institutions (Université Grenoble Alpes, CNRS, CEA, INRAE, ...), but also with key economic players in the area.
The Centre Inria de l'Université Grenoble Alpes is active in the fields of high-performance computing, verification and embedded systems, modeling of the environment at multiple levels, and data science and artificial intelligence. The center is a top-level scientific institute with an extensive network of international collaborations in Europe and the rest of the world.
Contexte et atouts du poste
Recent approaches in conversational AI for social assistive robots primarily rely on a human-robot substitution model, where the
robot replaces a human agent. The AnandaBot project (PEPR eNSEMBLE, France 2030, 2026-2030) instead investigates a human
(human/robot) approach, in which a robot sidekick accompanies and supports a human agent in a triadic collaborative task - much
as Ananda supported Buddha. The project is coordinated by Fabrice Lefèvre (LIA, Avignon Université) and brings together LIA
(Avignon Université), Inria (RobotLearn team), LISN (CNRS/Université Paris-Saclay), and AP-HP (Broca Hospital), with an applicationscenario in a gerontology day-hospital setting.
AnandaBot develops an audiovisual processing chain enabling triadic verbal and non-verbal interactions between a human user, a
human agent, and their robot sidekick. In this context, the offer is about continual audiovisual robot perception, and aims to develop
methods and algorithms to continuously extract cues about human behaviour from audio and visual data in real-world, socially
situated interactions, with two central requirements: (i) robustness to real-world perturbations, with quantitative estimation of the
reliability of extracted cues, and (ii) continuous adaptation to variations of these perturbations over time. We will cover three tasks:
behaviour understanding from visual inputs, continual audiovisual adaptation, and adaptation to simulated environments.
What do we offer?
- A 24-month postdoctoral contract at Inria Grenoble Rhône-Alpes, within the RobotLearn team.
- Remuneration according to Inria's postdoctoral salary scale, including standard Inria employee benefits (health insurance, paid leave, restaurant subsidy, etc.).
- Access to Inria's computing infrastructure and to the AnandaBot project's dedicated hardware (GPU servers, robot platforms).
- Integration in a well-funded, multi-site national consortium (LIA, Inria, LISN, AP-HP) with a clinical deployment site at Broca
Hospital.
- Support for travel to project meetings, conferences, and the yearly AnandaBot workshop.
References:
1. M. Marge, C. Espy-Wilson, and N. Ward, "Spoken Language Interaction with Robots: Research Issues and Recommendations," Report from the NSF Future Directions Workshop, 2019.
2. Y. Liu et al., "Continual learning for VLMs: A survey and taxonomy beyond forgetting," 2025.
3. Y. Xu, Y. Ban, G. Delorme, C. Gan, D. Rus, and X. Alameda-Pineda, "Transcenter: Transformers with dense representations for multiple-object tracking," IEEE TPAMI, 2022.
4. L. Vaquero, Y. Xu, X. Alameda-Pineda, V. M. Brea, and M. Mucientes, "Lost and found: Overcoming detector failures in online multi-object tracking," ECCV, 2024.
5. Y. Ban, X. Alameda-Pineda, L. Girin, and R. Horaud, "Variational Bayesian inference for audio-visual tracking of multiple speakers," IEEE TPAMI, 2019.
6. X. Alameda-Pineda et al., "Socially pertinent robots in gerontological healthcare," International Journal of Social Robotics, 2025.
7. A. Golmakani, M. Sadeghi, X. Alameda-Pineda, and R. Serizel, "A weighted-variance variational autoencoder model for speech enhancement," ICASSP, 2024.
8. J.-E. Ayilo, M. Sadeghi, R. Serizel, and X. Alameda-Pineda, "Diffusion-based unsupervised audio-visual speech enhancement," ICASSP, 2025.
9. S. Sadok, S. Leglaive, L. Girin, X. Alameda-Pineda, and R. Séguier, "A multimodal dynamical variational autoencoder for audiovisual speech representation learning," Neural Networks, 2024.
10. A. Ballou, X. Alameda-Pineda, and C. Reinke, "Variational meta reinforcement learning for social robotics," Applied Intelligence, 2023.
11. R. Aljundi, K. Kelchtermans, and T. Tuytelaars, "Task-free continual learning," CVPR, 2019.
12. S. Mo, W. Pian, and Y. Tian, "Class-incremental grouping network for continual audio-visual learning," ICCV, 2023.
Mission confiée
The postdoctoral researcher will contribute to the design, implementation, and evaluation of the audiovisual perception pipeline of
AnandaBot, under the supervision of Xavier Alameda-Pineda and Karteek Alahari. The position covers the following topics:
- Behaviour understanding from visual input: multi-person localisation and tracking, extraction of audiovisual cues (speech,
gaze, posture) relevant to socially situated triadic interactions, with generalisation to previously unseen people, objects, and
acoustic/visual conditions.
- Continual audiovisual adaptation: development of methods enabling the perception modules to adapt on-the-fly todistribution shifts (new speakers, new environments, sensor perturbations) without catastrophic forgetting, building on the team's prior work in continual and self-supervised learning.
- Adaptation to simulated environments: transfer of perception modules between the real-world experimental platform(deployed at Broca Hospital, AP-HP) and a simulation platform used for lower-cost, privacy-preserving training and evaluation,in coordination with the dialogue/engagement modules.
Principales activités
- Design and implement audiovisual perception algorithms (multi-person detection/tracking, speaker diarization, cue fusion)
robust to real-world social conditions.
- Develop continual-learning strategies for on-site, low-resource adaptation of audiovisual models, respecting the project's
privacy-by-design and data-efficiency requirements.
- Contribute to the simulation platform enabling transfer and evaluation of perception modules without requiring human
participants.
- Collaborate with AP-HP (Broca Hospital) on the definition of experimental protocols and the evaluation of perception modules
in laboratory and real-world (gerontology day-hospital) settings.
- Contribute to scientific publications, open-source releases, and the project's data/privacy compliance (informed consent,
secure on-site storage, CNIL and Ethics Committee submissions).
- Participate in project meetings (LIA, Inria, LISN, AP-HP) and the yearly AnandaBot workshop.
Compétences
- PhD in Computer Science, Machine Learning, Signal Processing, or a closely related field, completed or nearly completed at the
start date.
- Strong background in computer vision and/or audio(-visual) machine learning (e.g., multi-object tracking, speaker/source
localisation, multimodal fusion).
- Experience with deep learning frameworks (PyTorch or equivalent) and proficiency in Python.
- Familiarity with continual/lifelong learning and/or self-supervised learning is a strong plus.
- Experience with robotic platforms or human-robot interaction is a plus but not required.
- Good written and spoken English; French is not required but is a plus.
- Ability to work in a multidisciplinary, multi-site consortium.
Avantages
- Subsidized meals
- Partial reimbursement of public transport costs
- Leave: 7 weeks of annual leave + 10 extra days off due to RTT (statutory reduction in working hours) + possibility of exceptional leave (sick children, moving home, etc.)
- Possibility of teleworking and flexible organization of working hours
- Professional equipment available (videoconferencing, loan of computer equipment, etc.)
- Social, cultural and sports events and activities
- Access to vocational training
- Social security coverage under conditions
Rémunération
2788 € gross salary / month