Citation

If you find this model useful, please cite the following paper

@article{huang2024deciphering,
  title={Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate},
  author={Huang, Qidong and Dong, Xiaoyi and Zhang, Pan and Zang, Yuhang and Cao, Yuhang and Wang, Jiaqi and Lin, Dahua and Zhang, Weiming and Yu, Nenghai},
  journal={arXiv preprint arXiv:2410.07167},
  year={2024}
}

Downloads last month: 12

Safetensors

Model size

7.06B params

Tensor type

BF16

Inference Providers NEW

Image-Text-to-Text

This model is not currently available via any of the supported third-party Inference Providers, and the HF Inference API does not support transformers models with pipeline type image-text-to-text

Model tree for shikiw/LLaVA-v1.5-MoCa-7B

Base model

lmsys/vicuna-7b-v1.5

Finetuned

(50)

this model

shikiw
/

LLaVA-v1.5-MoCa-7B

Citation

Model tree for shikiw/LLaVA-v1.5-MoCa-7B

Dataset used to train shikiw/LLaVA-v1.5-MoCa-7B