VideoGemma

This repository contains code for VideoGemma multimodal language model.

VideoGemma combines LanguageBind video encoder with performant and flexible Gemma LLM in a LLaVA-style architecture.

Getting started

We recommend using Dev Containers to create the environment.

I don't want a container

Install PyTorch.

Install Python dependencies.

pip3 install -r requirements.txt

pip3 install git+https://github.com/facebookresearch/pytorchvideo.git@28fe037d212663c6a24f373b94cc5d478c8c1a1d

For checkpoint loading and model configuration see run_finetune.ipynb.

Pretrained checkpoints

Pretrained checkpoint for the model can be found here: HF 🤗.

The model's projector has been pretrained for 1 epoch on the Valley dataset.
LLM and the projector have been jointly fine-tuned using the Video-ChatGPT dataset.

Name		Name	Last commit message	Last commit date
Latest commit History 8 Commits
.devcontainer		.devcontainer
llava		llava
scripts		scripts
.DS_Store		.DS_Store
.gitignore		.gitignore
LICENSE		LICENSE
README.md		README.md
requirements.txt		requirements.txt
run_finetune.ipynb		run_finetune.ipynb

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Repository files navigation

VideoGemma

Getting started

I don't want a container

Pretrained checkpoints

About

Uh oh!

Releases

Packages

Uh oh!

Contributors 2

Uh oh!

Languages

License

tensorsense/videogemma

Folders and files

Latest commit

History

Repository files navigation

VideoGemma

Getting started

I don't want a container

Pretrained checkpoints

About

Resources

License

Uh oh!

Stars

Watchers

Forks

Releases

Packages 0

Uh oh!

Contributors 2

Uh oh!

Languages

Packages