Now showing 1 - 10 of 16
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    MAPEX: Modality-Aware Pruning of Experts for Remote Sensing Foundation Models
    Remote sensing data is commonly used in a wide range of tasks, such as natural disaster monitoring and land-use studies. For each task, scientists carefully choose appropriate modalities or leverage data from purpose-built instruments. Recent work on remote sensing foundation models pre-trains computer vision models on large amounts of remote sensing data to learn general-purpose representations. However, this progress comes at the cost of very large models, which are particularly challenging to deploy on edge devices due to their high inference costs. Moreover, downstream applications often rely on only a subset of modalities and operate under strict resource constraints, creating a mismatch between the large, multi-purpose models produced by pre-training and the lightweight, task-specific models needed in practice. We address this mismatch with MAPEX, a remote sensing foundation model based on mixture-of-modality experts. MAPEX is pre-trained on multi-modal remote sensing data using a novel modality-conditioned token routing mechanism that naturally encourages the emergence of modality-specialized experts. To apply the model on a specific task, we propose a modality-aware pruning technique that retains only the experts relevant to the task’s modalities, resulting in lightweight, task-specific models that can be directly extracted from the pre-trained foundation model, at no additional cost. Our approach yields efficient modality-specific models while simplifying fine-tuning and deployment for the modalities of interest. We experimentally validate MAPEX on diverse remote sensing datasets and show strong performance compared to fully supervised training and state-of-the-art remote sensing foundation models. Code will be available at https://github.com/HSG-AIML/MAPEX}{github.com/HSG-AIML/MAPEX.
    Type:
    Volume:
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    Physics-Guided Multitask Learning for Estimating Power Generation and CO 2 Emissions From Satellite Imagery
    (2023-05) ; ; ;
    IEEE Transactions on Geoscience and Remote Sensing (TGRS)
    Fossil fuel combustion produces large quantities of carbon dioxide (CO2), a major greenhouse gas (GHG), which is one of the main drivers of climate change. A quantitative assessment of GHG emissions is fundamental to predicting climate change effects, enforcing emission regulations, and monitoring pollution trading schemes. Unfortunately, the reporting of GHG emissions is only required in some countries, resulting in insufficient global coverage. At the same time, the transition from fossil fuels to zero carbon to limit climate change is at the heart of several ecological movements, hence the need for quantifying energy production, as well. In this work, we propose an end-to-end method to estimate power generation rates for fossil fuel power plants from satellite images, based on which we approximate GHG (CO2) emission rates. We present a physics-guided multitask deep-learning approach able to simultaneously predict from a single-satellite image of a power plant: 1) the pixel-area covered by plumes; 2) the type of fired fuel; and 3) the power generation rate. To ensure physically realistic predictions from our model we account for environmental conditions and empirical physical constraints. We then convert the predicted power generation rate into estimates for the rate at which CO2 is being emitted, using a fuel-dependent conversion factor. Experimental results show that our multitask learning approach improves the power generation estimation mean absolute error (MAE) by 23% compared to a single-task network trained on the same dataset.
    Type:
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    Know Your Attention Maps: Class-specific Token Masking for Weakly Supervised Semantic Segmentation
    Weakly Supervised Semantic Segmentation (WSSS) is a challenging problem that has been extensively studied in recent years. Traditional approaches often rely on external modules like Class Activation Maps to highlight regions of interest and generate pseudo segmentation masks. In this work, we propose an end-to-end method that directly utilizes the attention maps learned by a Vision Transformer (ViT) for WSSS. We propose training a sparse ViT with multiple [CLS] tokens (one for each class), using a random masking strategy to promote [CLS] token - class assignment. At inference time, we aggregate the different self-attention maps of each [CLS] token corresponding to the predicted labels to generate pseudo segmentation masks. Our proposed approach enhances the interpretability of self-attention maps and ensures accurate class assignments. Extensive experiments on two standard benchmarks and three specialized datasets demonstrate that our method generates accurate pseudo-masks, outperforming related works. Those pseudo-masks can be used to train a segmentation model which achieves results comparable to fully-supervised models, significantly reducing the need for fine-grained labeled data.
    Type:
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    Estimation of Power Generation and CO2 Emissions Using Satellite Imagery
    Burning fossil fuels produces large amounts of carbon dioxide (CO2), a major Greenhouse Gas (GHG) and a main driver of Climate Change. Quantification of GHG emissions related to power plants is crucial for accurate predictions of climate effects and for achieving a successful energy transition (from fossil-fuel to carbon-free energy). The reporting of such emissions is only required in some countries, resulting in insufficient global coverage. In this work, we propose an end-to-end method to predict power generation rates for fossil fuel power plants from satellite images based on which we estimate GHG emission rates. We present a multitask deep learning approach able to simultaneously predict: (i) the pixel-area covered by plumes from a single satellite image of a power plant, (ii) the type of fired fuel, and (iii) the power generation rate. To ensure physically realistic predictions from our model we account for environmental conditions. We then convert the predicted power generation rate into estimates for the rate at which CO2 is being emitted, using fuel-dependent conversion factors.
    Type:
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    A Multimodal Approach for Event Detection: Study of UK Lockdowns in the Year 2020.
    (IEEE Geoscience and Remote Sensing Society, 2022-07-19) ; ; ;
    Satellites allow spatially precise monitoring of the Earth, but provide only limited information on events of societal impact. Subjective societal impact, however, may be quantified at a high frequency by monitoring social media data. In this work, we propose a multi-modal data fusion framework to accurately identify periods of COVID-19-related lockdown in the United Kingdom using satellite observations (NO2 measurements from Sentinel-5P) and social media (textual content of tweets from Twitter) data. We show that the data fusion of the two modalities improves the event detection accuracy on a national level and for large cities such as London.
    Type:
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    Multitask Learning for Estimating Power Plant Greenhouse Gas Emissions from Satellite Imagery
    (Tackling Climate Change with Machine Learning workshop at NeurIPS., 2021-12-14) ; ; ;
    The burning of fossil fuels produces large amounts of carbon dioxide (CO2), a major Greenhouse Gas (GHG) and a main driver of Climate Change. Quantifying GHG emissions is crucial for accurate predictions of climate effects and to enforce emission trading schemes. The reporting of such emissions is only required in some countries, resulting in insufficient global coverage. In this work, we propose an end-to-end method to predict power generation rates for fossil fuel power plants from satellite images based on which we estimate GHG emission rates. We present a multitask deep learning approach able to simultaneously predict: (i) the pixel-area covered by plumes from a single satellite image of a power plant, (ii) the type of fired fuel, and (iii) the power generation rate. We then convert the predicted power generation rate into estimates for the rate at which CO2 is being emitted. Experimental results show that our model approach allows us to estimate the power generation rate of a power plant to within 139 MW (MAE, for a mean sample power plant capacity of 1177 MW) from a single satellite image and CO2 emission rates to within 311 t/h. This multitask learning approach improves the power generation estimation MAE by 39% compared to a baseline single-task network trained on the same dataset.
    Type:
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    Power Plant Classification from Remote Imaging with Deep Learning
    Satellite remote imaging enables the detailed study of land use patterns on a global scale. We investigate the possibility to improve the information content of traditional land use classification by identifying the nature of industrial sites from medium-resolution remote sensing images. In this work, we focus on classifying different types of power plants from Sentinel-2 imaging data. Using a ResNet-50 deep learning model, we are able to achieve a mean accuracy of 90.0% in distinguishing 10 different power plant types and a background class. Furthermore, we are able to identify the cooling mechanisms utilized in thermal power plants with a mean accuracy of 87.5%. Our results enable us to qualitatively investigate the energy mix from Sentinel-2 imaging data, and prove the feasibility to classify industrial sites on a global scale from freely available satellite imagery.
    Type:
    Scopus© Citations 5
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    GeoSANE: Learning Geospatial Representations from Models, Not Data
    Recent advances in remote sensing have led to an increase in the number of available foundation models; each trained on different modalities, datasets, and objectives, yet capturing only part of the vast geospatial knowledge landscape. While these models show strong results within their respective domains, their capabilities remain complementary rather than unified. Therefore, instead of choosing one model over another, we aim to combine their strengths into a single shared representation. We introduce GeoSANE, a geospatial model foundry that learns a unified neural representation from the weights of existing foundation models and task-specific models, able to generate novel neural networks weights on-demand. Given a target architecture, GeoSANE generates weights ready for finetuning for classification, segmentation, and detection tasks across multiple modalities. Models generated by GeoSANE consistently outperform their counterparts trained from scratch, match or surpass state-of-the-art remote sensing foundation models, and outperform models obtained through pruning or knowledge distillation when generating lightweight networks. Evaluations across ten diverse datasets and on GEO-Bench confirm its strong generalization capabilities. By shifting from pre-training to weight generation, GeoSANE introduces a new framework for unifying and transferring geospatial knowledge across models and tasks. Code is available at hsg-aiml.github.io/GeoSANE/
    Type:
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    Data-Efficient Deep Learning for Earth Observation
    Deep Learning methods have proven highly successful across a wide range of Earth Observation (EO)-related downstream tasks, such as image classification, image-based regression and semantic segmentation. Supervised learning of such tasks typically requires large amounts of labeled data, which oftentimes are expensive to acquire, especially for EO data. Recent advances in Deep Learning provide the means to drastically reduce the amount of labeled data needed to train models with a given performance and to improve the general performance of these models on a range of downstream tasks. As part of this tutorial, we will introduce and showcase the use of three such approaches that strongly leverage the multi-modal nature of EO data: Data Fusion, Multi-task learning and Self-supervised Learning. The fusion of multi-modal data may improve the performance of a model by providing additional information; the same applies to multi-task learning, which supports the model in generating richer latent representations of the data by means of learning different tasks. Self-supervised learning enables the learning of rich latent representations based on large amounts of unlabeled data, which are ubiquitous in EO, thereby improving the general performance of the model and reducing the amount of labeled data necessary to successfully learn a downstream task. We will introduce the theoretical concepts behind these approaches and provide hands-on tutorials for the participants utilizing Jupyter Notebooks. Participants, who are required to have some basic knowledge in Deep Learning with Pytorch, will learn through realistic use cases how to apply these approaches in their own research for different data modalities (Sentinel-1, Sentinel-2, land-cover data, elevation data, seasonal data, weather data, etc.). Finally, the tutorial will provide the opportunity to discuss the participants’ use cases.
    Type: