Next upHack for Humanity: San Francisco (powered by Google Gemini)
News

Researchers introduce VideoChat3, a 4B-parameter open video model, with weights still pending

A research team led by Nanjing University's MCG-NJU group introduced VideoChat3, a 4-billion-parameter open video model, in a July 17 paper; its weights remain unreleased.

D
Jul 17, 2026 · 1 min read

Researchers from Nanjing University’s MCG-NJU group, Shanghai AI Laboratory, Nanyang Technological University and Peking University introduced VideoChat3, a 4-billion-parameter open video language model, in a paper posted July 17.

VideoChat3 is a video multimodal large language model — a system that takes video and text as input — built for motion understanding, long-video reasoning, temporal grounding and live streaming. The team is pitching it as a smaller, cheaper alternative to larger proprietary video models.

At its core is an Inflated 3D Vision Transformer, or I3D-ViT, that compresses video 16 times across space and time, paired with an adaptive scheme that processes routine segments at 224-pixel resolution and reserves 448 pixels for critical moments. The design targets the main cost of video models: the flood of tokens that long clips generate.

The team says VideoChat3 surpasses comparably sized open models such as Qwen3-VL and Molmo2 on several video benchmarks while cutting inference latency by up to 54% on 2,048-frame inputs. Those figures are the authors’ own and have not been independently verified.

One caveat is central. The paper says the team is releasing model weights, training code and three datasets, but the project’s GitHub repository still lists the weights and data as pending — so the central open-source claim cannot yet be checked or reproduced by outside researchers.

Whether VideoChat3 lives up to “fully open” depends on that release landing.

More news