首页 > AI前沿 > VLM Fine-Tuning for End-to-End Combinatorial Optimization

VLM Fine-Tuning for End-to-End Combinatorial Optimization

arXiv自然语言 2026-09-29 18:00 5 阅读 查看原文

Large language models (LLMs) have provided a unified interface for end-to-end combinatorial optimization (CO), but textual serialization alone may obscure spatial and relational structures that are important for generating effective CO solutions.

This paper presents a general-purpose vision-language solver that augments textual instance descriptions with input-derived visual representations. A single vision-language model (VLM) is applied across different CO tasks and trained using supervised fine-tuning followed by verifier-guided reinforcement learning.

While the visual inputs contain no gold solutions or solution-derived information, our experiments show that the VLM generally improves solution quality over its text-only counterpart, with particularly clear gains on more complex CO problems such as CVRP and JSSP.

The advantage of visual information is more pronounced at large problem scales.