CIDR 2027 · Interactive artifact

One event query.
Three ways to execute it.

See how Concord substitutes a source-aligned transcript for video—or uses the transcript to restrict expensive video inference to a small fraction of the original video duration.

BaselineVideo → Localize
O1Transcript → Localize
O1 + O2 · Cross-modal rewrites
Example

Event localization

Loading example…

VRA Loading query…
Operator execution Ready to run
    Published evaluation · Table 2

    Less video, better localization

    Values below are frozen aggregate results over the evaluated lecture workload, not measurements recomputed from the trace above.

    Video processedO2 query
    Candidate recall
    MLLM costrelative to video-only
    F1 score
    Event-localization results.
    QueryVideoPrecisionRecallF1MLLM tokensCostTime
    Read the research

    Concord: A Video Relational Algebra for Cross-Modal Query Optimization

    Sultan Muratbek, Charisse Ivana Yeung, Chanwut Kittivorawong, and Alvin Cheung.

    MMDS repository ↗