Multi-Scale Solar Panel SegFormer (MiT-B2)

A Domain-Invariant Vision Transformer for Rooftop Solar Photovoltaic Detection & Segmentation from Satellite and Aerial Imagery.


Model Overview

  • Architecture: SegFormer with Hierarchical Mix Transformer (mit_b2) encoder and lightweight All-MLP decoder.
  • Parameters: 24.7 Million (~94 MB checkpoint).
  • Task: Binary Semantic Segmentation (0: Background / Roof, 1: Solar Photovoltaic Panel).
  • Input Resolution: 512 x 512 RGB imagery (supports arbitrary dimensions via patch sliding window).
  • Training Formulation: Boundary-Aware Compound Loss (0.4 BCE + 0.4 Soft Dice + 0.2 Laplacian Edge Loss).
  • Target Platforms: Google Maps / Google Earth satellite captures, sub-meter aerial orthomosaics, and high-resolution UAV drone surveys.

Key Capabilities & Advantages

  1. Multi-Scale Scale-Invariance (0.4x to 1.6x): Trained with multi-scale pyramid scaling. Detects panels across varying zoom levels on Google Maps and high-resolution aerial imagery.
  2. False-Positive Suppression (Hard Negatives): Ingests real Google Earth negative tiles (empty roofs, skylights, AC units, parking lots, roads) to suppress false alarms.
  3. Multi-Roof Architectural Generalization: Trained across 4 distinct roof categories:
    • Suburban Residential: Asphalt shingle and concrete tiles.
    • Commercial: Flat gravel and concrete roofs (shopping centers, warehouses).
    • Industrial: Corrugated metal and steel tile sheds.
    • European / Asian Heritage: Terracotta clay and brick roofs.

Quickstart: Running Inference in Python

import torch
import cv2
import numpy as np
import segmentation_models_pytorch as smp
from huggingface_hub import hf_hub_download

# 1. Download weights from Hugging Face
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
weights_path = hf_hub_download(
    repo_id="abhinav03700/solar-panel-segformer-mit-b2",
    filename="best_segformer_large.pth"
)

# 2. Instantiate SegFormer model
model = smp.Segformer(
    encoder_name="mit_b2",
    in_channels=3,
    classes=1,
    activation=None
).to(device)

model.load_state_dict(torch.load(weights_path, map_location=device))
model.eval()

# 3. Predict on an aerial / satellite image
def predict_solar(image_path, threshold=0.5):
    img = cv2.imread(image_path)
    img_rgb = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)
    h, w, _ = img_rgb.shape
    
    # Resize to 512x512
    img_resized = cv2.resize(img_rgb, (512, 512)).astype(np.float32) / 255.0
    tensor = torch.from_numpy(img_resized.transpose(2, 0, 1)).unsqueeze(0).to(device)
    
    with torch.no_grad():
        logits = model(tensor)
        probs = torch.sigmoid(logits).squeeze().cpu().numpy()
        
    mask = (probs > threshold).astype(np.uint8)
    mask_orig = cv2.resize(mask, (w, h), interpolation=cv2.INTER_NEAREST)
    return mask_orig

# Example:
# mask = predict_solar("rooftop_google_maps.png")

Citation

If you use this model or code in your research, please cite:

@article{solar_segformer_2026,
  title={Automated Rooftop Solar Photovoltaic Detection and Instance Segmentation from High-Resolution Aerial Orthomosaics},
  author={Abhinav et al.},
  journal={arXiv preprint},
  year={2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train abhinav03700/solar-panel-segformer-mit-b2