<?xml version='1.0' encoding='utf-8'?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>RS-Paper-Hub — VLM Papers</title>
  <id>https://rspaper.top/output/feed_vlm.xml</id>
  <link href="https://rspaper.top/output/feed_vlm.xml" rel="self" type="application/atom+xml" />
  <link href="https://rspaper.top" rel="alternate" type="text/html" />
  <updated>2026-09-15T05:51:08Z</updated>
  <subtitle>Latest remote sensing papers (last 7 days) — 10 entries</subtitle>
  <author>
    <name>RS-Paper-Hub</name>
    <uri>https://rspaper.top</uri>
  </author>
  <entry>
    <title>Transfer Learning for Socioeconomic Estimation in Forced-Displacement Settings</title>
    <link href="http://arxiv.org/abs/2609.15773v1" rel="alternate" type="text/html" />
    <id>http://arxiv.org/abs/2609.15773v1</id>
    <published>2026-09-14T00:00:00Z</published>
    <updated>2026-09-14T00:00:00Z</updated>
    <author>
      <name>Steven Ndung'u</name>
    </author>
    <author>
      <name>Adel Daoud</name>
    </author>
    <author>
      <name>Ismael Yacoubou Djima</name>
    </author>
    <author>
      <name>Hai-Anh H. Dang</name>
    </author>
    <author>
      <name>Patrick Michael Brock</name>
    </author>
    <summary type="text">Progress in inclusive household surveys has strengthened socioeconomic evidence for forcibly displaced populations, providing indispensable benchmarks on living conditions and welfare. However, these surveys remain resource-intensive and periodic, while conditions can change between rounds, particularly in settings affected by fragility, conflict, and violence. More frequently updated, spatially granular complementary evidence is therefore needed to identify where socioeconomic conditions may be changing between survey rounds and to inform operational prioritization. Earth observation and machine learning offer a scalable source of spatially explicit socioeconomic information. However, tools developed for general populations have not been systematically adapted and evaluated in forced displacement settings, where living conditions, settlement patterns, and displacement impacts may differ substantially. We address this gap by adapting a multimodal spatiotemporal vision transformer, pretrained on Demographic and Health Survey data from approximately 1.2 million households across 36 African countries, to forced displacement and host community settings in South Sudan, Cameroon, and Zambia. We develop and evaluate the updated, adapted model using socioeconomic indices derived from UNHCR FDS and RMS data. Our results show that satellite-derived geospatial covariates explain up to 66% of the variation in socioeconomic outcomes in camp-intersecting grids, with a mean absolute error (MAE) of 4.37 index points, and 41% in non-camp-intersecting areas, with an MAE of 5.41. The framework complements and adds value to periodic household surveys by filling critical spatial and temporal data gaps with regularly updated, model-based socioeconomic estimates. These estimates sustain insight between survey rounds and support timely humanitarian prioritization and field verification.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;Publication:&lt;/strong&gt; 13 pages, 7 figures&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Category:&lt;/strong&gt; Method&lt;/p&gt;</content>
    <category term="Machine Learning" />
    <category term="Artificial Intelligence" />
  </entry>
  <entry>
    <title>ANASSA: An Agentic AI Orchestration Framework for Spatial Intelligence</title>
    <link href="http://arxiv.org/abs/2609.14824v1" rel="alternate" type="text/html" />
    <id>http://arxiv.org/abs/2609.14824v1</id>
    <published>2026-09-13T00:00:00Z</published>
    <updated>2026-09-13T00:00:00Z</updated>
    <author>
      <name>Constantinos Papantoniou</name>
    </author>
    <author>
      <name>Brian Hilton</name>
    </author>
    <summary type="text">The emergence of large language models (LLMs) and large multimodal models (LMMs) has enabled a new class of agentic systems capable of integrating natural language understanding with tool-based execution. In geographic information systems (GIS), this shift is transforming traditional, expert-driven workflows into semiautonomous systems that can interpret user intent, construct spatial workflows, and execute geospatial analysis tasks. However, existing approaches remain limited by fragmented integration of reasoning, execution, and evaluation, particularly in complex, real-world environments. This study synthesizes recent advances in agentic GIS frameworks, benchmarks, and surveys to identify limitations in spatial reasoning, execution robustness, validation, governance, and evaluation. Building on these insights, it introduces ANASSA (Autonomous Neural Agents for Spatial Systems Architecture), an agentic AI orchestration framework that integrates structured spatial reasoning, multi-agent workflow orchestration, execution feedback, authoritative spatial validation, provenance, uncertainty handling, and human decision authority within a unified system design. The contribution is an architecture-level specification: eleven components across four layers, a six-step Geospatial AI Cognitive Loop, cross-component contracts, and governance mechanisms intended to make agentic geospatial workflows traceable, reproducible, and accountable. Empirical performance evaluation is reserved for implementation and deployment studies.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;Category:&lt;/strong&gt; Method&lt;/p&gt;</content>
    <category term="Artificial Intelligence" />
  </entry>
  <entry>
    <title>Speak to the City: Multimodal Resolution for Outside-the-Vehicle References</title>
    <link href="http://arxiv.org/abs/2609.14691v1" rel="alternate" type="text/html" />
    <id>http://arxiv.org/abs/2609.14691v1</id>
    <published>2026-09-13T00:00:00Z</published>
    <updated>2026-09-13T00:00:00Z</updated>
    <author>
      <name>Alireza Parchami</name>
    </author>
    <author>
      <name>Artin Saberpour</name>
    </author>
    <author>
      <name>Robin Connor Schramm</name>
    </author>
    <author>
      <name>Jürgen Steimle</name>
    </author>
    <author>
      <name>Ulrich Schwanecke</name>
    </author>
    <summary type="text">As autonomous vehicles and Extended Reality (XR) headsets enable novel in-car interactions, seamlessly querying physical landmarks, known as Outside-the-Vehicle Referencing (OVR), remains challenging due to ego-motion and referential ambiguity. We present a robust, multimodal OVR framework fusing user gaze and natural language to identify Points of Interest (POIs). To address the scarcity of dynamic vehicular data, we developed a VR-based pipeline synchronizing 360-degree transit videos with vehicle GNSS telemetry. Through a user study (N=46) mapping passenger head orientation into a 3D geospatial Digital Twin, we captured authentic gaze-speech behaviors. We subsequently trained a lightweight Transformer network, leveraging LLMs to dynamically align continuous spatial gaze vectors with discrete verbal context. Experimental results demonstrate high accuracy and low computational overhead, achieving an 83.33% Top-1 accuracy (87.72% Top-2) and an average inference time of 24.3 milliseconds. This real-time paradigm effectively resolves referential ambiguity, enabling context-aware spatial retrieval for passengers within the vehicle.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;Publication:&lt;/strong&gt; 11 pages, 7 figures&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Category:&lt;/strong&gt; Method&lt;/p&gt;</content>
    <category term="cs.HC" />
    <category term="cs.CL; Information Retrieval; Machine Learning" />
  </entry>
  <entry>
    <title>Towards foundation models for insurance risk modelling</title>
    <link href="http://arxiv.org/abs/2609.14576v1" rel="alternate" type="text/html" />
    <id>http://arxiv.org/abs/2609.14576v1</id>
    <published>2026-09-13T00:00:00Z</published>
    <updated>2026-09-13T00:00:00Z</updated>
    <author>
      <name>Christopher Blier-Wong</name>
    </author>
    <summary type="text">Claim narratives, images and sensor data contain information about insured risks that is difficult to use through existing actuarial models. Foundation models learn patterns from large datasets before being adapted to particular tasks. By turning these high-dimensional sources into variables or numerical representations, they could help insurers use more of the information they already collect, potentially reducing the experience needed to develop each application. For example, a language model could identify a worsening injury in a new claim note, allowing a reserving model to recognise the change in expected cost before the payments reveal the deterioration. In this paper, we review language, vision, geospatial, time series, tabular and scientific models, explaining existing insurance applications and potential future uses. Scientific models extend this approach to future weather and climate conditions: their simulations can inform loss estimates once local hazards are linked to asset damage, repair costs and insurance coverage. We propose a process to connect these model outputs to actuarial calculations and to assess their predictive contribution, stability and compliance with rules on information use. Evaluating these applications is difficult when final claim costs become known only after long delays, large losses are rare or patterns learned elsewhere fail to transfer to the target portfolio. Richer data can reveal private information and support finer risk classification, which can change access to insurance. Reusing the same models across insurers also creates dependence on shared predictions and providers.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;Category:&lt;/strong&gt; Method&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tasks:&lt;/strong&gt; CLS&lt;/p&gt;</content>
    <category term="q-fin.RM" />
  </entry>
  <entry>
    <title>Selective Tool Use for Agentic Change Visual Question Answering in Remote Sensing</title>
    <link href="http://arxiv.org/abs/2609.14523v1" rel="alternate" type="text/html" />
    <id>http://arxiv.org/abs/2609.14523v1</id>
    <published>2026-09-13T00:00:00Z</published>
    <updated>2026-09-13T00:00:00Z</updated>
    <author>
      <name>Yakoub Bazi</name>
    </author>
    <author>
      <name>Mohamad M. Al Rahhal</name>
    </author>
    <author>
      <name>Mohamed A. Mekhtiche</name>
    </author>
    <author>
      <name>Mansour Zuair</name>
    </author>
    <summary type="text">Change visual question answering (Change VQA) requires understanding semantic changes across bi-temporal remote sensing images. Although vision language models (VLMs) have shown promising performance on this task, they remain unreliable when answering questions that require explicit transition statistics, area measurements, or spatial information. To address this limitation, we propose a selective tool use framework in which a single VLM either answers directly or invokes a deterministic change analysis tool to obtain question specific evidence. Specifically, the selected tool operates on bi-temporal semantic maps and returns a structured observation, which the same VLM uses to generate its final answer. To support this framework, we construct a tool augmented extension of CDVQA covering eight question families and three tools for transition, spatial, and temporal analysis. Tool use supervision and observations are derived automatically from the original semantic annotations, without additional manual labeling. We then adapt Qwen3.5-4B using Low Rank Adaptation (LoRA) to jointly learn direct answering, tool invocation, and evidence conditioned answering. Experiments on 7,164 test questions show that selective tool use with reference semantic maps improves overall accuracy from 73.77% to 88.79% and average family accuracy from 69.11% to 89.65%. When the semantic maps are predicted automatically, the framework achieves 77.47% overall accuracy and 75.06% average family accuracy. These results demonstrate the benefit of question-specific semantic evidence for Change VQA, while highlighting the influence of semantic prediction quality on the resulting performance. Code and tool-augmented annotations will be made publicly available at https://github.com/yakoubbazi/ToolChangeVQA.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;Code:&lt;/strong&gt; &lt;a href="https://github.com/yakoubbazi/ToolChangeVQA"&gt;https://github.com/yakoubbazi/ToolChangeVQA&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Category:&lt;/strong&gt; Method&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tasks:&lt;/strong&gt; VQA;CD&lt;/p&gt;</content>
    <category term="Computer Vision" />
  </entry>
  <entry>
    <title>Physically Typed and Geometry-Aware Representations for Earth Foundation Models</title>
    <link href="http://arxiv.org/abs/2609.13868v1" rel="alternate" type="text/html" />
    <id>http://arxiv.org/abs/2609.13868v1</id>
    <published>2026-09-12T00:00:00Z</published>
    <updated>2026-09-12T00:00:00Z</updated>
    <author>
      <name>Rajiv Ranjan</name>
    </author>
    <summary type="text">Earth-observation (EO) foundation models have become exceptionally effective at learning se mantic, high-dimensional geospatial embeddings, while modern weather and climate models have demonstrated that Earth-specific geometry, spherical operators, meshes, and hybrid physical solvers can materially improve prediction. Yet these two advances are not equivalent. A conventional latent embedding has no inherent physical transformation law, whereas scalar fields, tangent polar-vector fields, axial/pseudovector quantities, covectors, and higher-order tensors transform differently under rotations, reflections, and changes of local coordinate frame. This proposal asks whether a general purpose Earth foundation model should preserve those distinctions explicitly, or whether standard embeddings plus augmentation already learn everything that matters. The central contribution is therefore not a more complicated architecture by assumption, but a staged falsification program. A compute-conscious ERA5 dry run first compares conventional, augmentation-matched, typed equivariant, and Hodge/Helmholtz variants under spatial, temporal, orientation, and low-data shifts. Only if explicit geometric typing yields reproducible improvements does the program advance toward a multimodal Earth foundation model in which semantic embeddings coexist with physically typed fields. The proposed gap is narrower and more defensible than claiming that current models ignore geometry entirely: several systems already respect spherical domain geometry, and emerging work explicitly learns scalar/vector fields on spheres. The unresolved question is whether foundation-scale, multimodal, parity-aware field typing produces practical gains beyond those existing approaches.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;Publication:&lt;/strong&gt; 18 pages&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Category:&lt;/strong&gt; Method&lt;/p&gt;</content>
    <category term="Computer Vision" />
    <category term="Machine Learning" />
  </entry>
  <entry>
    <title>Earth-Agent-Pro: Towards Real-World Full-Chain Earth Observation with Agents</title>
    <link href="http://arxiv.org/abs/2609.12533v1" rel="alternate" type="text/html" />
    <id>http://arxiv.org/abs/2609.12533v1</id>
    <published>2026-09-11T00:00:00Z</published>
    <updated>2026-09-11T00:00:00Z</updated>
    <author>
      <name>Zhutao Lv</name>
    </author>
    <author>
      <name>Chenhao Dang</name>
    </author>
    <author>
      <name>Yi Feng</name>
    </author>
    <author>
      <name>Yanpei Gong</name>
    </author>
    <author>
      <name>Xiaolei Wang</name>
    </author>
    <author>
      <name>Junyan Ye</name>
    </author>
    <author>
      <name>Conghui He</name>
    </author>
    <author>
      <name>Weijia Li</name>
    </author>
    <summary type="text">Real-world Earth observation (EO) agents must translate high-level scientific questions into executable workflows to acquire observations, prepare data, perform domain computations, and derive conclusions from runtime evidence. Existing EO agents typically start from supplied observations, while benchmarks typically provide prepared inputs or candidate answers, leaving full-chain open-world EO execution largely untested. We present Earth-Agent-Pro, an execution-adaptive Plan-and-Execute framework using expert-authored skills to constrain planning and runtime tool use. Workflow-centered structured memory records planned steps, accepted evidence, and their dependencies, enabling repair of only the affected workflow suffix when runtime evidence invalidates a step. Separate large language model adapters use sequence-level supervised fine-tuning for planner workflow composition and node-level group relative policy optimization with locally verifiable rewards for executor tool-argument grounding. Earth-Bench-Pro instantiates 248 expert-curated task cores as 744 questions under three matched regimes. Its 248 Open-World Execution questions span RGB imagery, spectral observations, and remote sensing products, pairing high-level requests with runtime data requirements, executable trajectories, and open-ended answers grounded in execution evidence. With a shared GPT-5 backbone, Earth-Agent-Pro achieves 66.13% LLM-as-Judge accuracy, exceeding ReAct by 20.95 points in this metric and 24.44 points in Tools-In-Order. Joint adapter tuning raises Qwen3.5-9B LLM-as-Judge accuracy from 38.31% to 50.00%, an 11.69-point gain over the untuned configuration. Planning-only evaluation and execution with the reference workflow show that the adapters improve workflow composition and argument grounding, respectively. Code and datasets will be released soon.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;Publication:&lt;/strong&gt; 18 pages, 8 figures&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Category:&lt;/strong&gt; Method&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tasks:&lt;/strong&gt; VG&lt;/p&gt;</content>
    <category term="Computer Vision" />
    <category term="Artificial Intelligence; cs.CL" />
  </entry>
  <entry>
    <title>Geospatial Foundation Models Capture Health-Relevant Dimensions of Place Beyond Conventional Social Risk Indices</title>
    <link href="http://arxiv.org/abs/2609.11689v1" rel="alternate" type="text/html" />
    <id>http://arxiv.org/abs/2609.11689v1</id>
    <published>2026-09-10T00:00:00Z</published>
    <updated>2026-09-10T00:00:00Z</updated>
    <author>
      <name>Nathaniel Hendrix</name>
    </author>
    <author>
      <name>Carl Y. Zhang</name>
    </author>
    <author>
      <name>Chris Heitzig</name>
    </author>
    <author>
      <name>Andrew Bazemore</name>
    </author>
    <author>
      <name>David H. Rehkopf</name>
    </author>
    <summary type="text">Area-based social risk indices summarize residents' socioeconomic conditions but incompletely capture physical features of place that may affect health. We evaluated whether numerical representations of physical place produced by four geospatial foundation model families from 2022 satellite data explained residual variance in tract-level associations between the Area Deprivation Index, Social Deprivation Index, and Social Vulnerability Index with health outcomes. We used LightGBM to predict variables from the American Community Survey and 40 chronic disease and health-behavior outcomes from CDC PLACES across 82,646 census tracts in the contiguous United States, evaluating performance across 10 held-out states. Among survey variables, models were moderately predictive of some variables including housing type (R-squared up to 0.54) but weak for disability, unemployment, and income disparity. For health outcomes, models explained up to 54% of variance left unexplained by social risk indices, with the largest gains for annual checkups, arthritis, and high blood pressure. Mean total variance explained by geospatial foundation models across the 40 health-related outcomes increased from 0.31 in the smallest tract-size decile to 0.39 in the largest. Geospatial foundation models capture health-relevant features of place not represented by conventional social risk indices and may usefully augment them in epidemiological analyses.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;Category:&lt;/strong&gt; Method&lt;/p&gt;</content>
    <category term="stat.AP" />
    <category term="Machine Learning" />
  </entry>
  <entry>
    <title>Geospatial AI, Dataverse Metadata, and the Study of Place-Based Government</title>
    <link href="http://arxiv.org/abs/2609.11674v1" rel="alternate" type="text/html" />
    <id>http://arxiv.org/abs/2609.11674v1</id>
    <published>2026-09-10T00:00:00Z</published>
    <updated>2026-09-10T00:00:00Z</updated>
    <author>
      <name>Danny EBanks</name>
    </author>
    <author>
      <name>Devika Jain</name>
    </author>
    <summary type="text">Harvard Dataverse hosts over 150,000 research datasets, but the geographic information those datasets carry is entered as free text by depositors and has never been assembled into a searchable structure. We construct a knowledge graph from the repository's public data and metadata, organizing 102,650 datasets within a 215,985-node network of 528,003 edges linking datasets to keywords, publications, subjects, journals, and locations. Of those datasets, 43,991 (42.9 percent) carry at least one geospatial field, geographic coverage, geographic unit, or a bounding box and 96.9 percent of all nodes sit in a single connected component, so datasets remain reachable from one another even when their geospatial metadata share nothing in common. A conservative keyword search identifies 7,654 geospatially tagged datasets (17.4 percent) as directly policy-relevant, with elections and legislatures the largest cluster, followed by government administration, health policy, transportation, and education. Five datasets illustrate how this metadata behaves across policy domains and spatial scales, and an extended use case shows how community language models, stance detection with geographic aggregation, and partisan language bridging tools can attach discourse to place. The central obstacle is place resolution: the same location appears as many disconnected nodes. We argue that the graph provides a concrete setting for developing AI-driven metadata enrichment and entity resolution, and we document its coverage skew toward American, city-level data.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;Category:&lt;/strong&gt; Method&lt;/p&gt;</content>
    <category term="Artificial Intelligence" />
  </entry>
  <entry>
    <title>From Pixels to Hierarchical Sequences: Quadtree Mask Encoding for Vision-Language Binary Change Detection</title>
    <link href="http://arxiv.org/abs/2609.09876v1" rel="alternate" type="text/html" />
    <id>http://arxiv.org/abs/2609.09876v1</id>
    <published>2026-09-09T00:00:00Z</published>
    <updated>2026-09-09T00:00:00Z</updated>
    <author>
      <name>Xiao An</name>
    </author>
    <author>
      <name>Ruikang Zhang</name>
    </author>
    <author>
      <name>Chen Zhong</name>
    </author>
    <author>
      <name>Xuli Shen</name>
    </author>
    <author>
      <name>Jiaxing Sun</name>
    </author>
    <author>
      <name>Jiang Wu</name>
    </author>
    <author>
      <name>Wei He</name>
    </author>
    <summary type="text">Dense change detection in remote sensing requires vision-language models (VLMs) to compare bi-temporal images and generate accurate pixel-level masks. Existing VLMs are largely confined to change captioning outputs, and the few that produce pixel-level masks still rely on external decoders or flat text-as-mask serialization, which are less effective for small and fragmented changes. We introduce QUAKE-CD, a framework that recasts dense change prediction as syntax-verifiable structured generation. QUAKE-CD represents binary change masks as grammar-constrained quadtree token sequences, making the masks compact, syntactically checkable, and deterministically decodable within an autoregressive generation space. We further construct QUAKE-CoT, which pairs these sequences with chain-of-thought traces grounded in visual evidence, and jointly optimizes textual reasoning and spatial dense prediction through a progressive curriculum followed by grammar-gated dual-reward RL. On QUAKE-CoT, QUAKE-CD achieves 78.31% accumulated F1, outperforming decoder-based and flat text-as-mask VLMs while producing more faithful bi-temporal reasoning.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;Publication:&lt;/strong&gt; 26 pages, 16 figures&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Category:&lt;/strong&gt; Method&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tasks:&lt;/strong&gt; IC;CD&lt;/p&gt;</content>
    <category term="Computer Vision" />
  </entry>
</feed>