TAN RUNJIAARMINE · ROBOTICS & EMBODIED AI
← ALL PROJECTS

CIS-RAM 2026 · First author

DR-TSLM

Understanding what people mean, across words, gestures and time.

Speech and co-speech gestures identify an object and its placement location.
Speech and co-speech gestures identify an object and its placement location.

The question

Natural instructions often leave objects and locations implicit. DR-TSLM aligns speech timestamps with video frames, then uses visual segmentation to resolve the intended object and destination.

The approach

A temporal-spatial language model connects deictic word detection, gesture context and spatial grounding. Its outputs include an object mask and target coordinates that downstream robot programs can use.

Evaluation

The manuscript reports 1,200 deictic expressions, 1,000 annotated scenes based on 100DOH, and 100 speech–gesture pairs from 10 participants. DR-TSLM-3B achieves 0.92 success on the temporal deictic reasoning task. This is a component-level result, not an end-to-end robot pick-and-place success rate.

Publication

CIS-RAM 2026

DR-TSLM: Deictic Reference Resolution Via Temporal-Spatial Large Language Model for Human-Robot Interaction

Runjia Tan, Shanhe Lou, Lan Yu, Xuesong Tian, Chen Lv