Automatic Camera-Based Inventory System
Imagine a warehouse with 10,000 SKUs—manual inventory every quarter halts shipments for three days. Discrepancies are discovered after the fact, leaving no time to correct them. We solve this with a computer vision system: cameras analyze shelves in real time, and YOLO algorithms (Ultralytics YOLOv8) detect each item. The system runs without stopping warehouse operations and achieves up to 98% accuracy, which is 2 times better than typical manual accuracy of 95%. For a 10,000-SKU warehouse, savings exceed 2 million rubles per year (approx. $24,000). This translates to annual savings of $24,000 in inventory labor costs. Our company has 5+ years of experience in warehouse automation and has completed 50+ projects, guaranteeing at least 95% accuracy from the first run.
How Camera Inventory Works
Fixed cameras above shelves take snapshots on a schedule (every 6–12 hours). Images pass through an object detection model—we use YOLOv8 (trained on 100+ product classes). The output is a list of SKUs with counts per shelf. These are reconciled with expected stock levels from your ERP. Discrepancies are flagged automatically, and the system can trigger replenishment orders. YOLOv8 processes a frame twice as fast as its predecessor, critical for hundreds of shelves. Optionally, we use NVIDIA Triton for batch inference, reducing p99 latency to 50ms.
Why 92–96% Accuracy Isn't Enough and How We Boost It
For most retailers, 95% accuracy seems acceptable, but at million-dollar turnover, each percentage point of discrepancy means direct losses. We push accuracy to 98% with three techniques: multi-angle capture, data augmentation, and reconciliation with POS data. Multi-angle capture reduces occlusion errors by 40%.
Achieving 98% Accuracy
Multi-angle capture is key. One camera cannot see items hidden behind others, so we install 2–3 cameras per aisle. For each frame, we apply perspective transformation to get a flat shelf view. Then we merge detections from different angles, removing duplicates by IoU (Intersection over Union). This reduces occlusion errors by 40%.
Reconciliation with the Accounting System
def reconcile(camera_counts: dict, system_counts: dict, tolerance_percent: float = 5.0) -> list[dict]: """Find discrepancies between physical and system counts""" discrepancies = [] all_skus = set(camera_counts) | set(system_counts) for sku in all_skus: camera_qty = camera_counts.get(sku, 0) system_qty = system_counts.get(sku, 0) if system_qty > 0: diff_pct = abs(camera_qty - system_qty) / system_qty * 100 else: diff_pct = 100 if camera_qty > 0 else 0 if diff_pct > tolerance_percent: discrepancies.append({ 'sku': sku, 'camera': camera_qty, 'system': system_qty, 'diff_percent': round(diff_pct, 1), 'severity': 'high' if diff_pct > 20 else 'medium' }) return sorted(discrepancies, key=lambda x: x['diff_percent'], reverse=True) Architecture & Tech Stack
The core of the system is a YOLO-based detector (PyTorch). Images are sent to a server with a GPU (NVIDIA T4 or A10); inference takes <100ms per image. We apply OpenCV perspective transformation to get a flat shelf view. For edge devices, we use TensorRT, speeding inference by 1.5× without accuracy loss.
class AutoInventorySystem: def __init__(self, detector_path: str, inventory_db_path: str): self.detector = YOLO(detector_path) self.db = InventoryDatabase(inventory_db_path) def run_inventory_cycle(self, shelf_images: dict) -> InventoryReport: """ shelf_images: {shelf_id: image} - photos of all shelves """ report = InventoryReport() for shelf_id, image in shelf_images.items(): shelf_counts = self._count_shelf(image, shelf_id) report.add_shelf(shelf_id, shelf_counts) # Compare with expected stock expected = self.db.get_expected_quantities() report.discrepancies = self._find_discrepancies( report.actual_counts, expected ) # Automatic update in ERP self.db.update_inventory(report.actual_counts) return report def _count_shelf(self, image: np.ndarray, shelf_id: str) -> dict: """Count products on one shelf""" detections = self.detector(image, conf=0.45) counts = {} for box in detections[0].boxes: sku = self.detector.model.names[int(box.cls)] counts[sku] = counts.get(sku, 0) + 1 return counts Perspective Distortion Handling
def create_shelf_rectified_view(image: np.ndarray, shelf_corners: list, output_size: tuple = (2000, 400)) -> np.ndarray: """ Flat (top-down) representation of the shelf for easier analysis shelf_corners: 4 corners of the shelf in the image """ pts_src = np.array(shelf_corners, dtype='float32') w, h = output_size pts_dst = np.array([ [0, 0], [w - 1, 0], [w - 1, h - 1], [0, h - 1] ], dtype='float32') M = cv2.getPerspectiveTransform(pts_src, pts_dst) rectified = cv2.warpPerspective(image, M, (w, h)) return rectified Drone-Based Inventory
class DroneInventoryController: def __init__(self, drone_api, inventory_system): self.drone = drone_api self.inventory = inventory_system self.waypoints = [] # pre-programmed shooting points async def run_inventory_mission(self) -> InventoryReport: images = {} await self.drone.takeoff() for waypoint in self.waypoints: await self.drone.fly_to(waypoint) await self.drone.stabilize(seconds=1.0) # Photos from multiple angles for better coverage for angle_offset in [0, -15, 15]: await self.drone.rotate(angle_offset) image = await self.drone.capture_image() images[f"{waypoint['shelf_id']}_{angle_offset}"] = image await self.drone.land() return self.inventory.run_inventory_cycle(images) ERP/WMS Integration
import requests class ERPIntegration: def __init__(self, erp_url: str, api_key: str): self.base_url = erp_url self.headers = {'Authorization': f'Bearer {api_key}'} def update_stock_levels(self, inventory: dict, location_id: str) -> dict: """Update stock levels in the ERP system""" stock_updates = [ { 'sku': sku, 'quantity': qty, 'location_id': location_id, 'source': 'camera_inventory', 'timestamp': get_iso_timestamp() } for sku, qty in inventory.items() ] response = requests.post( f'{self.base_url}/api/inventory/bulk-update', json={'updates': stock_updates}, headers=self.headers ) return response.json() | Inventory Type | Accuracy | Time |
|---|---|---|
| Fixed cameras (retail) | 92–96% | Continuous |
| Drone (1000 m² warehouse) | 90–95% | 20–40 min |
| Mobile robot | 94–98% | 30–60 min |
Implementation Process
- Analysis: we survey your warehouse, determine shelf count, angles, lighting. Select equipment.
- Design: we architect the system (cameras → server → ERP). Fine-tune the YOLO model for your product range.
- Implementation: mount cameras, deploy software, write integrations.
- Test: pilot run, reconcile with manual inventory, calibrate.
- Deploy: go live, train staff.
- Monitor & optimize: after launch, track accuracy, retrain the model for new SKUs as needed.
Deliverables
- Documentation: camera installation diagram, API spec, operator manual.
- Training: 2 days for warehouse team and 1 day for IT staff.
- Support: 12-month warranty on software, 3 months of free support.
- Metrics: dashboard showing accuracy, discrepancy count, inventory time.
Typical Mistakes and How We Avoid Them
- Occlusion: products hide each other. Solution: multi-angle capture.
- Changing lighting: we retrain the model on night-time frames.
- New SKUs: automatic registration via image search (embeddings).
| Scale | Timeline |
|---|---|
| Warehouse/store with fixed cameras | 6–9 weeks |
| Drone system + ERP | 10–16 weeks |
| Full autonomous system | 16–24 weeks |
Contact us for a free preliminary assessment of your warehouse. Get an individual estimate and pilot project—see the accuracy on real data. With our experience (5+ years, 50+ projects), we guarantee at least 95% accuracy from the first run. Time savings on inventory: up to 80%; loss reduction from discrepancies: up to 30%. Typical project cost starts from $50,000 for a warehouse with 10,000 SKUs.







