CIS-RAM 2026
CIS-RAM 2026 · First author
DR-TSLM
Understanding what people mean, across words, gestures and time.

The question
Natural instructions often leave objects and locations implicit. DR-TSLM aligns speech timestamps with video frames, then uses visual segmentation to resolve the intended object and destination.
The approach
A temporal-spatial language model connects deictic word detection, gesture context and spatial grounding. Its outputs include an object mask and target coordinates that downstream robot programs can use.
Evaluation
The manuscript reports 1,200 deictic expressions, 1,000 annotated scenes based on 100DOH, and 100 speech–gesture pairs from 10 participants. DR-TSLM-3B achieves 0.92 success on the temporal deictic reasoning task. This is a component-level result, not an end-to-end robot pick-and-place success rate.