首页 > AI前沿 > RS-Claw-Evolution: Environment-Feedback-Driven Evolution for Lightweight Remote Sensing Agents in Long-Horizon Tasks

RS-Claw-Evolution: Environment-Feedback-Driven Evolution for Lightweight Remote Sensing Agents in Long-Horizon Tasks

arXiv机器学习 2026-09-06 18:16 5 阅读 查看原文

Large language model-driven remote sensing (RS) agents offer a promising approach to automating geospatial analysis.

However, lightweight RS agents based on compact language models struggle with multi-step interactive tasks due to loss of long-horizon states, inefficient environmental feedback utilization, and sparse optimization signals.

We propose RS-Claw-Evolution

an environment-feedback-driven framework that progressively improves lightweight agents through three stages.

Interaction evolution

uses executable code to control observations, maintain intermediate states, and reduce context redundancy.

Experience evolution

combines failure-aware trajectory generation with error-turn masking to learn from informative failure-recovery experiences without imitating faulty actions.

Decision evolution

uses reinforcement learning with multi-dimensional environment rewards and turn-level advantage protection to optimize tool-use behaviors and improve credit assignment in long sequences.

On Earth-Bench,

the optimized Qwen3-4B-based agent achieves 65.9% accuracy in Autonomous Planning mode,

outperforming the untrained Qwen3-32B baseline (43.8%) and DeepSeek-V3.1 (60.8%),

while approaching GPT-5 (71.6%).

These results demonstrate that learning from environmental feedback can improve lightweight agents and narrow their performance gap with larger models in long-horizon RS tasks.