首页 > AI前沿 > VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model, and What Limits Its Visual Grounding

VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model, and What Limits Its Visual Grounding

arXiv自然语言 2026-08-09 12:46 6 阅读 查看原文

We build VectraYX-Vision-1B, a sub-2B Spanish/LATAM cybersecurity vision-language model for offline use, and measure what limits its visual grounding.

A frozen 1.04B-parameter decoder is coupled to a frozen vision encoder through a trainable projector: first SigLIP, then the Qwen2-VL-2B tower with a projector shaped for llama.cpp's mmproj export.

With SigLIP, after five fine-tuning defects were repaired, a nine-field extraction gate with a shuffled-image control passes the same 2/9 fields under every configuration that keeps the encoder and adds no text hint.

A frozen-feature probe explains why. Transplanting the Qwen2-VL tower with its own merger reads an 8-nibble address SigLIP never read (0.00 to 0.81 exact match).

On B8 (2,040 items, 34 fields, 16 templates, shuffled-image and best-constant controls), 9 fields pass on a Qwen-tower checkpoint, all within trained template-field pairs.

Screenshot tool identification (B6) sat at exactly 0.000. That was not a perception ceiling.

Stock Qwen2-VL-2B transcribes the same images at word recall 0.93 at our pixel budget (0.52 with our letterboxing).

Training only the projector on B6's task format and render family lifts B6 tool identification to 0.96 and recall to 0.48-0.51.

Perception of rendered text is therefore present and trainable in this frozen-backbone design.

The evidence is narrow. B6 is nearly in-distribution for that projector, which also forgets part of B8 (mean field accuracy 0.53 to 0.29 on a retention check); runs are single-seed, and no checkpoint yet has both.

We document four harness defects, retract an earlier B6 score, and release code, benchmarks and checkpoints.