Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher.
Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?
Some links on this page may take you to non-federal websites. Their policies may differ from this site.
-
Modern cities generate vast streams of urban dynamics data reflecting mobility demand, environmental conditions, and traffic patterns. The value of these data lies not only in individual modalities but in their integration—urban signals are highly interdependent, with changes in one modality often influencing others. Consequently, predicting any single urban dynamic requires information from multiple interrelated sources. Although numerous methods—ranging from deep learning models to recent LLM-based approaches—have been proposed, most are limited in scope. They either focus on single-modality prediction, rely on rigid model designs that lack flexibility, or overlook inter-modal dependencies. As a result, they struggle to adapt to dynamic urban conditions and suffer from degraded predictive performance across modalities. In this paper, we propose UniLLM, a unified large language model for multi-modal urban dynamics prediction. At its core, UniLLM introduces a Unified Cross-Modal Alignment Module that transforms heterogeneous urban data into latent representations while preserving modality-specific patterns and capturing cross-modal correlations through a contrastive learning objective. To support dynamic adaptation across tasks and modalities, we design a Routing-Aware Prompting Mechanism that learns soft prompts based on task context and modality semantics. Furthermore, a Multi-Modal Memory-Guided Adaptive Algorithm employs replay-based gradient coordination and Frank–Wolfe optimization to mitigate cross-modal catastrophic forgetting during fine-tuning. Extensive experiments across multiple cities and urban modalities demonstrate that UniLLM consistently outperforms state-of-the-art baselines. These results highlight UniLLM's potential as a flexible and robust forecasting model for real-world, multi-modal urban environments.more » « lessFree, publicly-accessible full text available April 20, 2027
-
The development of sensing technologies has broad-ened the scope of urban dynamics research. However, existing methods primarily focus on using isolated aspects of urban data, limiting their ability to capture the complex spatial-temporal dependency among different urban dynamics and generalizability across applications. Addressing these shortcomings requires a more comprehensive model capable of integrating multifaceted data, generating generalized representations adaptable to diverse applications and scenarios. In this paper, we introduce the Multifaceted SpAtial-TempoRal ContrastivE Learning framework, i.e., MARCEL, an innovative approach designed to learn robust, universally applicable, and adaptable representations of multifaceted urban dynamics through contrastive learning. MARCEL employs pretrained preliminary representation learning modules to extract distinct spatial-temporal dependencies inherent to each urban dynamic independently. It then features a Spatial-Temporal Contrastive Learning strategy to capture unified spatial-temporal patterns, including asynchronous, conflicting, and complementary behaviors across multifaceted urban dynamics. Additionally, MARCEL integrates a Multifaceted Knowledge Transfer mechanism to capture inter-dependencies among different urban dynamics and facilitate knowledge sharing. The learned representations are highly generalizable and can be applied effectively to various downstream tasks. Extensive experiments on real-world urban datasets demonstrate that MARCEL is effective and significantly outperforms state-of-the-art baselines.more » « lessFree, publicly-accessible full text available November 12, 2026
-
This paper introduces a novel approach to fine-tuning transformer-based models for various spatial-temporal downstream tasks. While fine-tuning approaches have shown remarkable success in fields like natural language processing, their efficacy in human-generated spatial-temporal data is often hindered by the complex spatial-temporal correlation and extensive reliance on labeled data. We introduce Knowledge Graph-Guided Spatial-Temporal Cross-task Fine-Tuning method, i.e., KG-STFT, a cross-task fine-tuning approach designed for adapting to data-scarce spatial-temporal tasks by leveraging rich knowledge embedded in similar downstream tasks and pre-trained model. Our KG-STFT framework utilizes i) a non-linear knowledge ensembler to capture and integrate knowledge embedded in different transformer blocks, and ii) constructs a task knowledge graph to "transfer" knowledge from data-rich to data-scarce tasks. Empirical experiments on real-world taxi trajectory data show that KG-STFT outperforms baselines, especially in data-scarce tasks, by leveraging task commonalities to improve fine-tuning.more » « lessFree, publicly-accessible full text available November 3, 2026
-
Given historical traffic distributions and associated urban conditions observed in a city, the conditional urban traffic estimation problem aims at estimating realistic future projections of the traffic under a set of new urban conditions, e.g., new bus routes, rainfall intensity, and travel demands. The problem is important in reducing traffic congestion, improving public transportation efficiency, and facilitating urban planning. However, solving this problem is challenging due to the strong spatial dependencies of traffic patterns and the complex relations between the traffic and urban conditions. Recently, we proposed a Complex-Condition-Controlled Generative Adversarial Network C3-GAN, which tackles both of the challenges and solves the urban traffic estimation problem under various complex conditions by adding a fixed embedding network and an inference network on top of the standard conditional GAN model. The randomly chosen embedding network transforms the complex conditions to latent vectors, and the inference network enhances the connections between the embedded vectors and the traffic data. However, a randomly chosen embedding network cannot always successfully extract features of complex urban conditions, which indicates C3-GAN is unable to uniquely map different urban conditions to proper latent distributions. Thus, C3-GAN would fail in certain traffic estimation tasks. Besides, C3-GAN is hard to train due to vanishing gradients and mode collapse problems. To address these issues, in this article, we extend our prior work by introducing a new deep generative model, namely, C3-GAN+, which significantly improves the estimation performance and model stability. C3-GAN+ has new objective, architecture, and training algorithm. The new objective applies Wasserstein loss to the conditional generation case to encourage stable training. Shared convolutional layers between the discriminator and the inference network help to capture spatial dependencies of traffic more efficiently, part of the shared convolutional layers are used to update the embedding network periodically aiming to encourage good representation and avoid model divergence. Extensive experiments on real-world datasets demonstrate that our C3-GAN+ produces high-quality traffic estimations and outperforms state-of-the-art baseline methods.more » « less
An official website of the United States government

Full Text Available