Concord

A Video Relational Algebra for Cross-Modal Query Optimization

Sultan Muratbek*, Charisse Ivana Yeung*, Chanwut (Mick) Kittivorawong, Alvin Cheung

*Equal contribution · UC Berkeley

Semantic video queries let users analyze video content by issuing natural-language prompts to a multimodal large language model (MLLM). Such queries are increasingly popular, but their expressiveness comes at a steep cost. An MLLM may process hours of media to return only seconds of relevant output, making naive execution slow, expensive, and inaccurate.

We propose Concord, a system for expressing and optimizing semantic video queries. Concord makes three contributions. First, we introduce Video Relational Algebra (VRA), a nested algebra over videos, transcripts, frames, and object tracks that captures common semantic video operations. Second, we derive approximate VRA rewrites that can reduce MLLM usage and improve result quality. For narrated video, Concord either processes transcripts instead of video or uses them to identify video clips for MLLM processing. For cross-camera queries without narration, detection and tracking replace whole-video semantic association performed by an MLLM with a track-level relational join.

Third, we evaluate Concord on real-world videos. Across 4.59 hours of soccer broadcasts and 3.92 hours of lectures, transcript-to-video queries send only 5.32% and 2.47% of source-video duration to the MLLM and reduce MLLM cost by 85.0% and 87.2%, respectively. On two five-second highway clips containing 18 manually adjudicated cross-camera vehicles, a Detect–Track–Join query improves F1 from .364 to .813 while making no MLLM calls.

@misc{muratbek2026concord,
      title  = {Concord: A Video Relational Algebra for Cross-Modal Query Optimization},
      author = {Sultan Muratbek and Charisse Ivana Yeung and
                Chanwut Kittivorawong and Alvin Cheung},
      year   = {2026},
}