Muhammad Raza - Computer Vision & AI Agent | YOLO - MediaPipe - LLM Real-Time

Computer Vision AI Agent
YOLO – MediaPipe – LLM Real-Time

Overview

Most machine learning projects fail between the prototype and production. I've shipped 47+ that didn't.

You have a working concept — or a clear problem involving cameras, video, or image data. The challenge is making it fast, accurate, and stable under real-world conditions. Wrong framework choices. Inference too slow for live video. Models that break the moment lighting, angle, or environment changes. And systems that detect things but can't reason about them or act on them autonomously. That's exactly where most builds stall.

I design and build real-time computer vision pipelines that go all the way — from model training to live deployment — and increasingly, from visual perception to autonomous AI agents that understand, decide, and narrate.

Object detection · Machine learning · Pose estimation · Multi-camera tracking · Segmentation · Re-identification · Anomaly detection · OCR & ANPR · Optical flow · Depth estimation · LLM-powered reasoning · Agentic decision pipelines

While most CV engineers stop at training the model, I go further:
→ Accelerated inference with TensorRT, ONNX, OpenVINO, and FP16/INT8 quantization (up to 5× faster)
→ LLM agents layered over CV pipelines for real-time decisions, alerts, and natural language outputs
→ Mobile deployment via CoreML (iOS) and TFLite (Android) with 10+ live apps shipped
→ Edge deployment on Jetson, OpenVINO, Apple Neural Engine, and CUDA/cuDNN
→ End-to-end pipeline: camera input → training → optimization → real-time actionable output

Key Accomplishments:

  • ⭐ Generated $5M+ in client revenue
  • ⭐ Delivered 100+ end-to-end computer vision systems
  • ⭐ Successfully launched 2 SaaS products
  • ⭐ Real-time sports AI for 7+ sports, improving analytics for 15+ teams
  • ⭐ Mobile AI on iOS (Core ML) & Android (TFLite), powering 10+ apps
  • ⭐ Medical imaging AI for 5+ hospitals: tumor detection, ultrasound, test strips
  • ⭐ Model optimization: up to 5× faster inference using FP16/INT8, ONNX, TensorRT, OpenVINO
  • ⭐ Agentic CV systems that perceive, reason, and act without human input

If your project involves cameras, video, or images — and you need it fast, accurate, fully deployed, and intelligent enough to reason and act autonomously — I am the engineer you are looking for.

🎯 YOLO Detection 🧍 Pose Estimation 🏋️ Sports AI 🛒 Retail AI 🛡️ CCTV Analytics 🔄 Tracking 🧠 ML Pipelines 🤖 AI Agents 💬 LLM Integration

Book a consultation

Development & IT Consultation
30 min Zoom meeting
AI & Machine Learning AI Data Annotation & Labeling AI Integration Mobile App Development
Book a consultation

What I Build

  • Real-time object detection, tracking & pose estimation pipelines
  • Sports AI — action recognition, scoring, player analytics
  • iOS & mobile on-device inference (Core ML, Vision Framework, Swift)
  • YOLO model training, fine-tuning & conversion (TensorRT, ONNX, CoreML)
  • LLM-powered AI agents with visual perception & decision pipelines
  • Edge AI optimization — FP16/INT8, OpenVINO, TensorRT acceleration
  • OCR/ANPR, multi-camera tracking, PPE & safety monitoring systems

Client Reviews

Frequently
Asked Questions

I design and build real-time computer vision pipelines from model training to live deployment — and increasingly, from visual perception to autonomous AI agents that understand, decide, and narrate. Services include object detection, pose estimation, multi-camera tracking, OCR/ANPR, LLM-powered reasoning, and agentic decision pipelines.

YOLO, MediaPipe, OpenCV, Core ML, Vision Framework, TensorFlow, PyTorch, TensorRT, ONNX, OpenVINO, Swift/SwiftUI, Python, Flutter, FastAPI, and AWS. Specialized in accelerated on-device inference with FP16/INT8 optimization for mobile and edge devices.

51+ jobs on Upwork including iOS Tennis AI, ChefAI Smart Kitchen Monitoring, Volleyball AI Vision Tracking, AI Fitness Rep Counter, Flutter YOLO v8 Object Detection App, PPE Factory Safety Monitoring, real-time sports tracking/scoring systems, YOLO model training & conversion for mobile/desktop/web, and edge AI optimization projects.

Absolutely. Clients say: "He did a fantastic job. He's reliable, responsible, and highly recommended, especially for computer vision problems." I integrate as an embedded CV/AI engineer, handling everything from architecture to deployment — whether joining an existing team or working solo end-to-end.

Yes — iOS is a core specialty. I build fully on-device CV apps in Swift/SwiftUI using Apple's Vision Framework and Core ML, targeting iOS 17+ on A12 Bionic or newer. Zero cloud latency, complete privacy. Projects include real-time tennis tracking, pose-based fitness apps, and action recognition systems.