A drone survey can show an entire landscape in extraordinary detail.
It can reveal roads, tracks, drainage lines, construction activity, bare ground, vegetation and, sometimes, the scattered signatures of waste dumping. But there is a catch. A detailed orthomosaic is not really an image in the everyday sense. It is a georeferenced map built from many aerial photographs, potentially tens of thousands of pixels wide and high.
In this case, a single DroneDeploy GeoTIFF can be approximately 52,000 by 52,000 pixels.
That is far too large to hand to a standard object-detection model in one piece.
So the question became: how can we teach a model to find likely waste piles across a landscape-sized drone map, while retaining the geographic coordinates needed for real monitoring, reporting and field response?
The answer is a tile-then-detect workflow built around YOLOv7.
A model that never sees the whole map
The model does not inspect the full orthomosaic.
Instead, the system divides the huge drone map into fixed-size RGB tiles, each normally 2,500 by 2,500 pixels. Think of it as placing a grid over a very large aerial map, then giving the computer one square at a time.
The model itself then works on an even smaller, standardised version of each tile. Before inference, each 2,500-pixel crop is letterboxed and resized to 640 by 640 pixels, the input size used by YOLOv7.
This makes the task computationally manageable, but it also means the model is looking at a downsampled view of the original drone imagery. In effect, it operates at roughly four times lower spatial detail than the native orthomosaic.
That trade-off is deliberate. The model needs enough context to recognise the shape, colour and texture of a probable waste pile, but the whole workflow also needs to be fast enough to process a large survey area.
Turning mapped waste into training data
A detection model is only as useful as the examples used to teach it.
The starting point here is a set of waste-pile labels provided as GeoJSON polygons. These are geographic features, drawn in the same coordinate reference system as the drone orthomosaic. They represent real-world mapped locations, not ordinary image annotations.
Before YOLOv7 can learn from them, those mapped polygons must be translated into the pixel coordinates of the raster.
The pipeline reads the GeoTIFF’s coordinate transformation, converts each label from eastings and northings into raster row and column positions, and identifies the image tiles that intersect each labelled pile. It then clips each bounding box to the tile and converts it to YOLO’s compact annotation format. All coordinates are normalised relative to the individual tile. There is only one target category.
This is a deliberately focused design. The model is not trying to separate concrete, timber, plastic, excavated soil and mixed construction waste into multiple categories. It is trained to answer a simpler operational question:
Does this image tile contain something that visually resembles a dumped pile?
Only tiles containing at least one labelled pile are retained for the training dataset. Optional image augmentation introduces variations such as flips, transposition, hue and brightness shifts, and 90-degree rotations. These transformations help reduce the risk that the model learns irrelevant features, such as a particular orientation of the drone flight path or a specific lighting condition.
The final dataset is divided into training and validation subsets, with an 80/20 split.
Transfer learning gives the model a head start
The detector uses YOLOv7, a high-performance object-detection architecture. Rather than learning from scratch, it starts from the official pre-trained YOLOv7 weights and is fine-tuned for the single pile class.
This process is called transfer learning.
The original model has already learned broad visual features from a very large image dataset: edges, shapes, textures, patterns and object-like structures. The fine-tuning stage adapts those general capabilities to a much narrower environmental question, namely, whether an RGB image contains a likely pile of dumped material.
The training configuration uses 640 by 640-pixel inputs, a batch size of 16, a planned 70 training epochs and the default stochastic gradient descent optimiser. Mosaic augmentation remains enabled, allowing the detector to encounter more varied arrangements of image content during training.
The aim is not to create a magical all-seeing waste detector. It is to develop a model that can rapidly screen a large drone survey and return a manageable set of likely locations for review.
The important part happens after detection
In a conventional computer-vision project, a bounding box might be the final result.
Here, it is only the halfway point.
YOLOv7 returns a prediction in the pixel coordinates of a tile. To make that prediction useful in environmental monitoring, the workflow must add the tile’s position within the full orthomosaic, reconstruct the bounding box in full-mosaic pixel coordinates, and then transform it back into map coordinates.
The conversion follows the affine transformation stored in the GeoTIFF. In simple terms, the workflow translates an image-space box into a rectangle with a real-world location.
The output is a GeoJSON FeatureCollection containing predicted pile locations. That is what makes the model operationally relevant. The output can be loaded into a GIS, overlaid on roads, drainage channels, development areas or protected sites, compared with older surveys, reviewed by an analyst and used to plan a site visit.
The model detects visual objects. The geospatial workflow turns them into places.
From a box on a tile to a decision on the ground
The immediate use case is not automated enforcement. A detected box means that the model has found a location whose visual characteristics resemble the labelled examples of piles or dumping. It is a candidate for review, not proof that a site is illegal, recently deposited or environmentally harmful.
A sensible operational workflow would look like this:
- Fly the site and generate a georeferenced orthomosaic.
- Divide the mosaic into image tiles and run the trained YOLOv7 model across each tile.
- Convert predicted boxes back into geographic coordinates.
- Publish the detections as a GIS layer or web map.
- Review high-confidence detections alongside recent imagery, known construction areas, access routes and previous inspection records.
- Prioritise field inspection where there is a likely new deposit, a large accumulation, a sensitive location or repeated activity.
- Feed verified findings, including false positives and missed piles, back into the training dataset.
This is where the system starts to become environmental intelligence rather than a one-off AI demonstration.
The longer-term ambition aligns with earlier dumping-monitoring work in Riyadh, where remote sensing was intended to support rapid detection, notification, reporting and more efficient allocation of site visits.
The current limitations matter
A useful environmental AI system should state its limitations plainly.
The current pipeline uses non-overlapping tiles. That creates an obvious edge case: a pile located on the boundary between two crops may be split, partly visible or missed entirely. Overlapping tiles, followed by cross-tile non-maximum suppression, would be a logical improvement in a future version.
The model also works only from RGB imagery. It does not use elevation, near-infrared information, thermal data, contextual land-use layers or time-series change detection. These additional data sources could help distinguish a genuine dumping event from visually similar features such as active construction, stockpiled materials, exposed ground or earthworks.
There is also a data-export issue to resolve. The current GeoJSON writer rebuilds each feature dictionary several times, which may result in malformed features. The export function should be corrected and tested in a GIS before any operational deployment.
Finally, and most importantly, the detector identifies a visual target class. It does not determine ownership, waste composition, volume, date of deposition, permit status or legal liability.
Those questions still require human assessment, site context and, where necessary, field verification.
What comes next
The next phase is about testing the model in the real diversity of landscapes it will encounter.
Can it recognise piles on pale soil and dark ground? Does it work equally well near roads, construction sites and drainage lines? How does it handle shadows, partial vegetation cover, different drone flight conditions and piles that are small, irregular or spread across tile boundaries?
The answers will shape the next version of the dataset and the pipeline.
Possible improvements include overlapping tiles, a corrected GeoJSON export routine, polygon segmentation rather than rectangular bounding boxes, more representative negative examples, multi-date analysis and the addition of contextual GIS layers. The resulting system could also feed a dashboard that highlights new, persistent or expanding sites for inspection.
A drone orthomosaic is far too large for a detector to process at a glance. But by breaking it into tiles, teaching YOLOv7 to recognise a single meaningful target, and carefully reconnecting predictions to geographic space, it becomes possible to turn aerial imagery into a practical map of places that deserve attention.
That is the adventure in the data jungle: not just finding objects in images, but turning pixels into environmental decisions.
Technical
Model: YOLOv7, single detection class, “pile”Training source: RGB crops from a georeferenced drone orthomosaicTraining labels: GeoJSON polygons transformed from raster CRS coordinates to pixel coordinatesTile size: 2,500 × 2,500 pixels, non-overlappingModel input: 640 × 640 pixelsTraining split: 80% training, 20% validationInference thresholds: confidence 0.50, IoU 0.45Output: predicted bounding boxes reprojected to GeoJSON polygons
Suggested visual sequence
- Hero image: orthomosaic view of a real pile, overlaid with one restrained YOLO detection box and the title.
- Scale image: full orthomosaic overview beside one highlighted 2,500-pixel tile.
- Label-transfer diagram: GeoJSON polygon in map coordinates → pixel coordinates → clipped YOLO label within a tile.
- Pipeline graphic: GeoTIFF → tiles → YOLOv7 → tile detections → affine reprojection → GeoJSON map layer.
- Honesty panel: examples of a true positive, a false positive and an edge-of-tile failure.
- Final image: detections shown in GIS context with roads, wadi lines, development areas or environmental receptors.
The best final headline may actually be “YOLO Sees Piles, GIS Knows Where They Are”. It captures the central technical insight in a way that is memorable, technically honest and appropriate for the exploratory, applied tone of Data Jungle Adventures.
Photo by Daniel Miksha on Unsplash
Leave a Reply