LLM-jp-4-VL is a series of vision-language models developed by LLM-jp.
This repository provides sample code for running inference with the LLM-jp-4-VL models.
Install dependencies:
uv syncSee cookbooks/basic.py for a runnable inference example covering text-only, single-image, multi-image, and multi-turn inputs. It works with both llm-jp/llm-jp-4-vl-9b (reasoning) and llm-jp/llm-jp-4-vl-9B-beta (non-reasoning); switch models by editing model_id at the top of the file.
To reproduce the evaluation results reported in our blog post, please refer to simple-evals-mm, our VLM evaluation framework.
This code is released under the Apache 2.0 license.
If you find our work useful, please consider citing the following papers:
@misc{sugiura2026jaglebuildinglargescalejapanese,
title={Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision-Language Models},
author={Issa Sugiura and Keito Sasagawa and Keisuke Nakao and Koki Maeda and Ziqi Yin and Zhishen Yang and Shuhei Kurita and Yusuke Oda and Ryoko Tokuhisa and Daisuke Kawahara and Naoaki Okazaki},
year={2026},
eprint={2604.02048},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2604.02048},
}
@misc{sugiura2026jammevalrefinedcollectionjapanese,
title={JAMMEval: A Refined Collection of Japanese Benchmarks for Reliable VLM Evaluation},
author={Issa Sugiura and Koki Maeda and Shuhei Kurita and Yusuke Oda and Daisuke Kawahara and Naoaki Okazaki},
year={2026},
eprint={2604.00909},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2604.00909},
}