CogVideo
CogVideoX (2024) and CogVideo — text-and-image-to-video models from Tsinghua/Zhipu AI.
TL;DR · 30-second scan
CogVideo (Python) — CogVideoX (2024) and CogVideo — text-and-image-to-video models from Tsinghua/Zhipu AI.
You want a research-backed open video model with a permissive license.
AI Video Generation
The most academically respected open video model. CogVideoX (the current version) produces cleaner outputs than most competitors at 6-second clips. The 5B parameter version fits on a 24GB GPU, the 12B needs 48GB+. Documentation is dense but the code is clean. Apache license and Chinese academic origin means it will keep getting free updates. Not as glossy as Higgsfield but a solid open foundation to fine-tune on.
You want a research-backed open video model with a permissive license.
You want plug-and-play — expect a couple of hours of setup on your first run.
Add this badge to your README to show your project is curated on StackPicks. Free, lightweight (180×28 SVG), and gives your visitors a one-click way to see honest take + alternatives.
[](https://stackpicks.dev/repo/cogvideo)
<a href="https://stackpicks.dev/repo/cogvideo"><img src="https://stackpicks.dev/api/badge/cogvideo" alt="Featured on StackPicks" width="180" height="28" /></a>
Are you the maintainer of zai-org/CogVideo? Add the badge and we'll feature your project in the weekly curator newsletter.