In densely built cities, point sensors give scattered readings. Managers need a concentration map with a 100 m grid and an accurate 48-hour forecast. Deterministic dispersion models require dozens of parameters and often err by 50–70% in variable wind. ML models trained on historical data reduce the error to 15–20%. Over 5 years, we have completed 20+ AI-monitoring projects in cities with different climates and building densities.
System Architecture
Data → Processing → Storage → Analytics → Visualization Data: ├── Government posts (Rosgidromet, FBU CLM) ├── IoT sensors (own/partner) ├── Satellite (Sentinel-5P, MODIS) ├── Mobile stations (cars, bicycles) └── NWP weather forecasts (Rosgidromet API, Open-Meteo) Processing: ├── Low-cost sensor (LCS) calibration ├── Quality assurance/quality control (QA/QC) ├── Spatial interpolation └── Air quality forecasting Storage: └── TimescaleDB (temporal) + PostGIS (spatial) Analytics: ├── AQI calculation ├── Trend analysis ├── Source attribution └── Health impact estimation Visualization: └── Web portal + mobile app | Component | Technology | Purpose |
|---|---|---|
| Data collection | MQTT, LoRaWAN | Receiving data from sensors and APIs |
| Storage | TimescaleDB + PostGIS | Time series + spatial data |
| ML models | XGBoost, LSTM, U-Net | Interpolation, forecasting, source attribution |
| Visualization | Leaflet, React Native | City map and mobile app |
How ML Interpolation Works
Stations are always fewer than needed. For a 100 m air quality map, interpolation is necessary. Kriging is 30% less accurate than ML interpolation, especially in areas with local sources (factories, roads). We use XGBoost with spatial covariates: distance to highways, NDVI, building density.
| Method | Accuracy (RMSE) | Resolution | Data Requirements |
|---|---|---|---|
| Kriging | 25–30 µg/m³ | Depends on network | Only stations |
| XGBoost | 10–15 µg/m³ | 100 m | Stations + covariates |
| U-Net | 8–12 µg/m³ | 30–100 m | Satellite + stations |
def spatial_air_quality_model(station_readings, spatial_covariates): """ Train on station_readings Predict for the entire urban 100×100 m grid """ X = pd.merge(station_readings, spatial_covariates, on=['lat', 'lon']) # Spatial features X['distance_to_highway'] = ... X['distance_to_industry'] = ... X['ndvi'] = ... # greening X['building_density'] = ... # building density model = XGBRegressor().fit(X, X['pm25']) return model # Predict for the entire city grid grid = create_city_grid(city_boundary, resolution=100) grid['predicted_pm25'] = model.predict(grid[feature_cols]) Deep Learning for spatial mapping: U-Net with multispectral satellite imagery + station readings → PM2.5 map at 30–100 m resolution. Training on simultaneous station and satellite data. Sensor maintenance cost savings — up to 40%.
Why LSTM Outperforms Traditional Models
Key forecast factors:
- Meteorology: wind (speed and direction determine transport), atmospheric stability (mixing height), precipitation (PM washout)
- Sources: industrial emissions, transport, heating
- Photochemistry: O3 formation and secondary particles (PM2.5)
LSTM + Weather Attention model:
class AirQualityForecastModel(nn.Module): def __init__(self): self.pollutant_encoder = LSTM(n_pollutants, 64) self.weather_encoder = LSTM(n_weather_vars, 64) self.cross_attention = CrossAttention(64, 64) self.decoder = nn.Linear(128, n_pollutants * forecast_hours) Horizon: 24/48/72 hours. Achievable MAPE: < 15% for 48-hour PM2.5 forecast.
Air Quality Index (AQI)
AQI calculation per "Method for Calculating the Atmospheric Pollution Index" (RD 52.04.667):
def calculate_aki(concentrations: dict) -> float: """ AQI = Σ (C_i / MPC_daily_i) for i pollutants AQI < 5 — standard air quality """ aqi = 0 for pollutant, conc in concentrations.items(): mpc = MPC_DAILY[pollutant] aqi += conc / mpc return aqi Color coding:
- Green: AQI < 5 (normal)
- Yellow: 5–7 (slight pollution)
- Orange: 7–14 (moderate)
- Red: > 14 (high, health hazard)
More about Air Quality Index.
Mobile App for Citizens
Features:
- Current AQI at user location
- City air quality map
- AQI forecast for 24/48 hours
- Recommendations: safe to walk/exercise?
- Notifications when thresholds exceed
Personalized recommendations:
- Asthmatics/allergy sufferers: stricter notification threshold
- Cyclists: optimal time/route considering AQI
- Parents with children: playground quality index
Source Attribution
Positive Matrix Factorization (PMF) decomposes PM2.5 chemical composition into sources: industry, transport, residential heating, natural (sea salt, dust). We use EPA PMF 5.0 — the official tool for receptor modeling.
from scipy.optimize import nnls # G = F × C (observations = sources × contributions) # PMF minimizes the weighted sum of squared residuals # under non-negativity of F and C Result: "30% of city PM2.5 from metallurgy, 40% from transport, 20% from residential heating." This forms the basis for regulatory decisions.
How to Assess ROI of the Monitoring System
Direct savings: reduced fines for exceeding norms (up to 2 million rubles per year in industrial zones), lower laboratory measurement costs (replaced by ML interpolation), optimized filter operation based on actual load. Indirect effect — increased citizen trust and city investment attractiveness.
Example of low-cost sensor calibration
We use gradient boosting to correct drift: the model is trained on pairs of cheap sensor readings and a reference station, considering temperature and humidity. This reduces error from 50% to 10%.What's Included in the Work
- Data audit: assess quality and completeness of existing measurements, choose optimal stack.
- Architecture design: data flow diagram, model selection, prototyping.
- Sensor calibration: adjust low-cost sensors against reference stations.
- ML model development: interpolation, forecasting, source attribution.
- Web portal and mobile app creation: map, AQI, alerts.
- Testing and validation: comparison with independent stations, MAPE < 15%.
- Documentation and training: handover of code, API, instructions for ecologists.
- 3-month support: monitoring, model retraining, updates.
Timeline: basic IoT network + AQI calculation + map + mobile app — 8–10 weeks. System with ML forecasting, source attribution, and regulatory API — 4–5 months.
Work stages:
- Data audit — 1 week
- Design — 1–2 weeks
- Sensor calibration — 2–3 weeks
- ML model development — 3–4 weeks
- Interface development — 3–4 weeks
- Testing — 2–3 weeks
- Documentation and training — 1–2 weeks
Contact us — we'll analyze your data in 2 days. Get a consultation on stack selection and work scope estimation. We'll evaluate your project in 2 business days. We guarantee quality thanks to certified engineers and 5+ years of experience in AI monitoring.







