<?xml version='1.0' encoding='utf-8'?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>RS-Paper-Hub — Agent Papers</title>
  <id>https://rspaper.top/output/feed_agent.xml</id>
  <link href="https://rspaper.top/output/feed_agent.xml" rel="self" type="application/atom+xml" />
  <link href="https://rspaper.top" rel="alternate" type="text/html" />
  <updated>2026-08-13T00:53:54Z</updated>
  <subtitle>Latest remote sensing papers (last 7 days) — 2 entries</subtitle>
  <author>
    <name>RS-Paper-Hub</name>
    <uri>https://rspaper.top</uri>
  </author>
  <entry>
    <title>GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning</title>
    <link href="http://arxiv.org/abs/2608.10494v1" rel="alternate" type="text/html" />
    <id>http://arxiv.org/abs/2608.10494v1</id>
    <published>2026-08-11T00:00:00Z</published>
    <updated>2026-08-11T00:00:00Z</updated>
    <author>
      <name>Xin Xiao</name>
    </author>
    <author>
      <name>Jiang Zhong</name>
    </author>
    <author>
      <name>Junnan Zhu</name>
    </author>
    <author>
      <name>Yingchao Feng</name>
    </author>
    <author>
      <name>Peijin Wang</name>
    </author>
    <author>
      <name>Yidan Zhang</name>
    </author>
    <author>
      <name>Kaiwen Wei</name>
    </author>
    <summary type="text">Earth observation (EO) agents construct scientifically valid tool workflows and ground their conclusions in current geospatial evidence. This is challenging because EO workflows are constrained by sensing semantics, product dependencies, spatial and temporal compatibility, and parameter requirements. Existing agents often search a broad operation space for each query, while recent self-evolving systems do not fully organize heterogeneous EO trajectories into reusable knowledge across different decision levels. To solve this problem, we present GeoForge, a training-free, self-evolving framework that transforms completed trajectories into a structured nonparametric execution state. GeoForge constrains the operation space according to the sensing context, then retrieves a task-conditioned prior from three complementary memories. Workflow Graph Memory captures global operation order, Action-Level Experiences provide local corrections, and the Adapted Skill Standard Operating Procedure preserves procedural and data constraints. The retrieved prior guides tool execution, while current observations remain the basis of the final answer. After each task, a safety-gated distillation process converts grounded trajectories into reusable execution knowledge for future retrieval. This execution, distillation, and reuse loop improves planning without updating the backbone LLM. Experiments on multiple geospatial benchmarks demonstrate that GeoForge consistently improves both task accuracy and tool-use trajectory quality across diverse LLM backbones, while substantially reducing tool-planning and reasoning errors for most LLMs.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;Category:&lt;/strong&gt; Method&lt;/p&gt;</content>
    <category term="Artificial Intelligence" />
    <category term="cs.MA" />
  </entry>
  <entry>
    <title>LMM Modality Transfer: A Pre-requisite for Autonomous GIS Agents</title>
    <link href="http://arxiv.org/abs/2608.06948v1" rel="alternate" type="text/html" />
    <id>http://arxiv.org/abs/2608.06948v1</id>
    <published>2026-08-07T00:00:00Z</published>
    <updated>2026-08-07T00:00:00Z</updated>
    <author>
      <name>Ivan Majic</name>
    </author>
    <author>
      <name>Zexian Huang</name>
    </author>
    <author>
      <name>Franziska Hübl</name>
    </author>
    <author>
      <name>Krzysztof Janowicz</name>
    </author>
    <author>
      <name>Meilin Shi</name>
    </author>
    <author>
      <name>Mina Karimi</name>
    </author>
    <author>
      <name>Zilong Liu</name>
    </author>
    <author>
      <name>Alexandra Fortacz-Lazan</name>
    </author>
    <summary type="text">AI models are becoming increasingly adept at understanding and processing spatial information, thereby facilitating agentic problem-solving in spatial tasks and workflows. However, most of the research on their spatial capabilities (e.g., spatial reasoning) has focused on the textual modality as input and output. This contrasts with the human approach to GIS workflows, where text and visual modalities are often used together, interchangeably, and in a complementary manner. Thus, to truly achieve an automated GIS analysis pipeline or carry out human-designed GIS workflows, AI models --- Large Multimodal Models (LMMs) in particular --- need to be able to seamlessly transition between image- and text-based modalities that are traditionally used in such workflows. We present a modality transfer task that (1) asks an LMM to first describe an input image of colored squares in a regular grid, and (2) asks a new LMM instance to re-generate an image of the original spatial scene using the textual description output by the former model. This task quantifies the ability of LMMs to transfer spatial information between image and text modalities. Ultimately, by examining the modality transfer capability of LMMs through the lens of spatial information theory, this work highlights a critical bottleneck: achieving strong and robust geospatial understanding in LMMs requires rigorous, multi-modal alignment. Our results indicate that recent LMMs (here from OpenAI) still struggle with modality transfer, when tasked with re-generating an image of a simple spatial grid of color squares.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;Category:&lt;/strong&gt; Method&lt;/p&gt;</content>
    <category term="Artificial Intelligence" />
  </entry>
</feed>