What I have been reading →

Arjun Rao

Arjun Rao

I am a Computer Science PhD student at the University of British Columbia (UBC), affiliated with the Vector Institute and the Pioneer Centre for AI. I am broadly interested in representation learning for spatial and geographic applications.

Before graduate school, I worked as a Machine Learning Scientist at Orbital Sidekick in San Francisco and interned at the NASA Jet Propulsion Laboratory. I graduated from The Chinese University of Hong Kong majoring in Financial Technology and Computer Science.

Service is a major component of my non-research endeavors. Since 2024, I have worked with the Prison Mathematics Project in the capacity of both a student instructor and an advocate during pre-clemency instructional sessions. I also volunteer at Boulder County Jail's math circle as a volunteer tutor. Since 2021, I have been affiliated with the 414LIFE program to end youth violence in Milwaukee, Wisconsin.

arjun.rao[at]ubc.ca

Selected Publications

Arjun Rao, Sebastian Loeschcke, Anthony Fuller, Isaac Corley, Nico Lang✰Equal advising, Evan Shelhamer✰Equal advising, "Planetary Feature Fields are Scalable Earth Representations"
Preprint, 2026.
[Paper]

Satellite observations, precomputed embeddings, and map products describe the same evolving Earth, yet are stored as independent, petabyte-scale data products. Their continued growth calls for compact representations of multiple products while preserving spatial and temporal detail. We introduce Planetary Feature Fields (PFFs), which exploit redundancy across data products by modeling them jointly as continuous functions of space and time at planetary scale. PFFs are spatially local explicit-implicit (hybrid) neural fields. Each field shares a factored feature volume—a decomposition of an explicit 3D grid with smaller factors—across products, while lightweight implicit decoders reconstruct individual products across multiple timesteps. PFFs reconstruct EO products over space and time more accurately than single-product fields at matched compression rates. At 1800× compression relative to the uncompressed source data, reconstructed features retain approximately 90% or more of the performance achieved with the original features on pixel-level segmentation, change detection, and patch-level classification tasks. PFFs can add new timesteps by extending their factored feature volumes and add new products by attaching new decoders, while leaving existing outputs unchanged. PFFs reduce end-to-end feature access latency by an order of magnitude relative to evaluated API and cloud-storage pipelines.

Arjun Rao, Ruth Crasto, Tessa Ooms, David Rolnick, Konstantin Klemmer, Marc Rußwurm, "Localized, High-resolution Geographic Representations with Slepian Functions"
ICML 2026.
[Paper]

Geographic data is fundamentally local. Disease outbreaks cluster in population centers, ecological patterns emerge along coastlines, and economic activity concentrates within country borders. Machine learning models that encode geographic location, however, distribute representational capacity uniformly across the globe, struggling at the fine-grained resolutions that localized applications require. We propose a geographic location encoder built from spherical Slepian functions that concentrate representational capacity inside a region-of-interest and scale to high resolutions without extensive computational demands. For settings requiring global context, we present a hybrid Slepian-Spherical Harmonic encoder that efficiently bridges the tradeoff between local-global performance, while retaining desirable properties such as pole-safety and spherical-surface-distance preservation. Across five tasks spanning classification, regression, and image-augmented prediction, Slepian encodings outperform baselines and retain performance advantages across a wide range of neural network architectures.

Arjun Rao, Marc Rußwurm, Konstantin Klemmer, Esther Rolf, "Measuring the Intrinsic Dimension of Earth Representations"
ICLR 2026.
[Paper]

Within the context of representation learning for Earth observation, geographic Implicit Neural Representations (INRs) embed low-dimensional location inputs (longitude, latitude) into high-dimensional embeddings, through models trained on geo-referenced satellite, image or text data. Despite the common aim of geographic INRs to distill Earth's data into compact, learning-friendly representations, we lack an understanding of how much information is contained in these Earth representations, and where that information is concentrated. The intrinsic dimension of a dataset measures the number of degrees of freedom required to capture its local variability, regardless of the ambient high-dimensional space in which it is embedded. This work provides the first study of the intrinsic dimensionality of geographic INRs. Analyzing INRs with ambient dimension between 256 and 512, we find that their intrinsic dimensions fall roughly between 2 and 10 and are sensitive to changing spatial resolution and input modalities during INR pre-training. Furthermore, we show that the intrinsic dimension of a geographic INR correlates with downstream task performance and can capture spatial artifacts, facilitating model evaluation and diagnostics. More broadly, our work offers an architecture-agnostic, label-free metric of information content that can enable unsupervised evaluation, model selection, and pre-training design across INRs.

Arjun Rao, Esther Rolf, "Using Multiple Input Modalities Can Improve Data-Efficiency and OOD Generalization for ML with Satellite Imagery", TerraBytes Workshop @ ICML 2025, Proceedings of Machine Learning Research (PMLR). Spotlight Best Poster Award
Preliminary version accepted to ML4RS Workshop @ ICLR 2025. Oral Best Student Paper Award
[Paper]

A large variety of geospatial data layers is available around the world ranging from remotely-sensed raster data like satellite imagery, digital elevation models, predicted land cover maps, and human-annotated data, to data derived from environmental sensors such as air temperature or wind speed data. A large majority of machine learning models trained on satellite imagery (SatML), however, are designed primarily for optical input modalities such as multi-spectral satellite imagery. To better understand the value of using other input modalities alongside optical imagery in supervised learning settings, we generate augmented versions of SatML benchmark tasks by appending additional geographic data layers to datasets spanning classification, regression, and segmentation. Using these augmented datasets, we find that fusing additional geographic inputs with optical imagery can significantly improve SatML model performance. Benefits are largest in settings where labeled data are limited and in geographic out-of-sample settings, suggesting that multi-modal inputs may be especially valuable for data-efficiency and out-of-sample performance of SatML models. Surprisingly, we find that hard-coded fusion strategies outperform learned variants, with interesting implications for future work.

Arjun Rao, Steffen Mauceri, Andrew K Thorpe, Jake Lee, Brian Bue, Siraput Jongaramrungruang, Riley Duren, "Improving Imaging Spectrometer Methane Plume Detection with Large Eddy Simulations", American Geophysical Union, Fall Meeting 2021 [Website]

Methane is a highly potent greenhouse gas, and accurately measuring methane emissions is an important step in addressing global environmental change. A large number of human-caused methane emissions originate from small ground structures known as "point-source emitters". These structures typically span a few meters wide and release highly concentrated methane. Current work on detecting these emissions use machine learning models that are trained to identify point-source emissions from aerial imagery. However, these models require large amounts of expensive labelled data, and have shown an inability to adapt to new terrain and environments. To increase the amount of methane emission data available to scientists, we use a collection of synthetic methane measurements. This data is generated with a mathematical model built to realistically replicate the three-dimensional distribution of methane observed in aerial imagery. We train a Convolutional Neural Network on the generated synthetic data, and test our model on realistic datasets containing airborne imagery from diverse terrain and environmental conditions. Training models on a combination of synthetic and real methane data helps reduce false positives on previously unseen scenes.

Research Experience

University of British Columbia / Vector Institute
Department of Computer Science, UBC
Computer Science PhD Student
Present
Orbital Sidekick
Machine Learning Engineer
Jun 2022 – Aug 2024
NASA Jet Propulsion Laboratory
Research Intern, Caltech SURF@JPL
Jun 2021 – Aug 2021