Client Delivery • 3 Months
VLM-Based Fall Detection
Benchmarking vision-language models to catch falls in video — built for a real client engagement
Python, GPT-4o, Claude, Gemini, Qwen2.5-VL
A desktop app that analyzes video for fall events using vision-language models, built with a small team over a 3-month client engagement. Rather than betting on one model, it benchmarks GPT-4o, Claude, Gemini, and Qwen2.5-VL side by side so the client could actually see the trade-offs before committing to one.

01 / CONTEXT
What needed to change.
The key decision was not to declare a single vision-language model the winner. The system had to make trade-offs visible for a real client engagement, because model behaviour, latency, and video reasoning quality change depending on the provider and the clip.
02 / WHAT I BUILT
A system built around the real constraints.
The desktop application presents video fall-event analysis through a simple interface and supports Qwen2.5-VL by default alongside compatible OpenAI, Anthropic, and Google Gemini models. The configuration layer exposes each model’s runtime settings, while environment variables keep provider keys isolated from the app. The workflow lets the client compare GPT-4o, Claude, Gemini, and Qwen2.5-VL before committing to a model path.
03 / WHERE IT STANDS
What works now, and what’s still being tuned.
The project runs locally with Python 3.13.9 and model-provider credentials configured in the environment. It remains an evaluation tool rather than a finished operational deployment: output quality depends on the selected model and input video, and the next step is validating the benchmark workflow with more representative fall scenarios.