Computer Vision Engineer at Wobot.ai
Shipping real-time vision pipelines that run at the edge.
I build and maintain production video-analytics systems — NVIDIA DeepStream and Hailo pipeline servers, Triton inference deployments, and the tooling that keeps them observable. IEEE-published and patent-filed on AeroVision, my work on aerial computer vision.
Who I Am
I'm Shreyansh Rao, a computer vision engineer who works on the unglamorous half of AI: the part where a model has to run on real hardware, on real camera feeds, without falling over at 3 a.m.
At Wobot.ai I work across the video-inference stack — migrating detection services onto the DeepStream pipeline server, restructuring the Hailo pipeline server, chasing root causes on edge devices in production, and upgrading Triton Inference Server across environments. Alongside that I build the tooling that makes the stack legible: automated Grafana log analysis, exponential log throttling, and FPS/latency reporting that turns a wall of logs into a few numbers you can act on.
My research work became AeroVision, published at IEEE and filed as a patent. Before Wobot I worked on cloud-deployed LLM applications at Trinity Packaging Digital and AI workflow automation research at Coding Jr.
Outside of work I read papers, break my own side projects, and try to make things run faster than they did yesterday.
DeepStream & Hailo pipeline servers, multi-stream RTSP, edge deployment
Triton Inference Server, TensorRT, model compatibility & migration
RCA on live systems, Grafana log automation, throttling and metrics
IEEE-published and patent-filed on AeroVision
Where I Work
Wobot.ai turns ordinary camera feeds into operational intelligence for enterprises. I work on the inference layer that makes that possible — the pipeline servers, edge runtimes and observability tooling that have to hold up across thousands of live streams. Most of my work sits at the seam between models and infrastructure: keeping DeepStream, Hailo and Triton stable while new detection services get folded in.
My Journey
Working across the production video-inference stack: migrating services onto the DeepStream pipeline server, restructuring the Hailo pipeline server, running RCAs on edge devices, and upgrading Triton Inference Server. Also built the log-analysis and throttling tooling the team uses to debug live pipelines.
Deployed AI-powered applications to the cloud using Docker. Built a cloud-hosted platform for the ASB Audit Risk Analyzer, integrated an Ollama Phi-3 LLM for intelligent responses, and applied prompt engineering to improve output accuracy — strong exposure to production deployments, containerization and scalable system design.
Analysed productivity bottlenecks at companies like ShareChat and CRED, researched their tech stacks, and recommended AI-powered automation opportunities within their development pipelines. Also contributed to Planto AI Copilot, automating software workflows such as code generation, testing and deployment.
Publication & IP
Jamming-Resilient Radar-Based Aerial Object Classification Using Temporal Attention Networks
Radar keeps working in rain, fog and darkness where cameras and LiDAR struggle, which is exactly why autonomous systems lean on it — but a deep-learning classifier trained on clean radar returns can fall apart the moment someone jams the signal. This paper proposes a classification architecture built to survive that: a shared CNN extracts features from a sequence of Range-Doppler frames, a temporal attention network learns to weight the cleaner frames more heavily than the corrupted ones, and a second, auxiliary head is trained in parallel to detect whether jamming is present at all. Training both heads together through one composite loss pushes the network toward features that hold up under interference rather than just fitting the clean case.
What I've Built
A full-stack computer vision app that processes live video streams with YOLOv8 and renders detections in an interactive web dashboard. Supports multiple concurrent camera feeds, custom model fine-tuning and async inference workers.
Real-time player tracking and re-identification using YOLOv8 and DeepSORT. Players are matched across multiple video feeds via appearance embeddings and cosine similarity, so each person keeps one consistent global ID as they move between cameras.
A chatbot that lets you upload several PDFs at once and ask questions across all of them. Documents are chunked, embedded and retrieved semantically, so answers stay grounded in the source text with citations back to the originating document.
A fine-tuned BERT sentiment model for noisy social-media text, with a preprocessing pipeline that handles emoji, slang and code-mixed input. Served behind a REST API with batched inference.
Automated attendance using dlib and face_recognition. Encodes faces in real time, matches against a registered database, and exposes a web interface for management plus CSV export of attendance records.
A hybrid recommender combining content-based similarity over movie metadata with collaborative signals from user ratings. Vectorized descriptions, genres and cast are compared by cosine similarity, served through a Streamlit interface with poster fetching.
Technical Stack
= daily-driver tools across current work & projects
Academic Background
Specialization in Artificial Intelligence and Machine Learning. Coursework in Deep Learning, Computer Vision, Data Structures & Algorithms, Probability & Statistics, Linear Algebra and Software Engineering. Final-year research became the AeroVision paper, published at IEEE and filed as a patent.
Recognition
Published on IEEE Xplore for research on deep-learning-based aerial computer vision
The method behind the AeroVision paper has been filed as a patent
Built the Grafana log-analysis automation and log-throttling systems now used to debug live pipelines
Upgraded Triton Inference Server 25.03 → 26.03 across E2E and staging with full model compatibility
A full overview of my experience, skills, education and projects in a clean, single-page format.
📄 Download PDFGet In Touch
I'm always up for talking about production computer vision, edge inference, or interesting AI/ML problems in general. If you're hiring, collaborating, or just curious about the AeroVision work — reach out.