Computer Visionwith VLMs
An introduction to vision-language models for builders. Understand how modern AI sees images — and ship your first vision pipeline in 5 hours.
How a VLM answers
Image → patches
The image is split into a grid of tokens
Vision encoder
Patches become embeddings the LLM can read
“The invoice total is $1,240.”
Language model reasons over both
What You'll Walk Away With
Short and dense by design. Built for engineers and technical PMs who want to add vision to their products without a research detour.
A working mental model
Understand what happens between an image going in and an answer coming out — so failures stop being mysterious.
Practical prompting skills
Extraction, grounding, comparison, and multi-image workflows you can apply the same day at work.
One real pipeline
Leave with a document-understanding pipeline you built yourself, ready to adapt to your own data.
The 5-Hour Curriculum
Five focused modules you complete at your own pace. Every module ends with something you built.
From Pixels to Meaning
45 minWhy classic computer vision hit a wall, and how vision-language models changed what "seeing" means for software. CLIP, contrastive learning, and the core intuition you need.
How a VLM Actually Sees
60 minImage patches, vision encoders, and projection into the language space. No math beyond what you need — just enough to predict when a model will fail.
Prompting for Vision Tasks
75 minHands-on: visual question answering, grounding and bounding boxes, structured extraction from images, and multi-image reasoning. Working with GPT-4o, Claude, and Qwen-VL.
Real Applications
60 minOCR-free document parsing, UI and screenshot understanding, product catalog tagging, and video frame analysis. Build one working pipeline end-to-end.
Limits, Evals & Model Choice
60 minHallucinated details, counting failures, spatial reasoning gaps. How to eval a VLM for your use case and pick between hosted APIs and open weights.
Join the Waitlist
Be the first to know when the course launches in June 2026.
Start whenever you want — learn at your own pace
Questions? Contact us at hello@theagentcamp.com
Frequently Asked Questions
Do I need machine learning experience?
No. We explain the internals at the level of intuition, not equations. If you can write Python and call an API, you have everything you need.
Which models will we use?
We work hands-on with hosted APIs (GPT-4o, Claude) and one open-weights model (Qwen-VL) so you see both sides of the trade-off.
Is 5 hours really enough?
For an introduction, yes. The goal is a solid mental model and one working pipeline — not covering every architecture. You leave knowing exactly what to learn next.
How does self-paced work?
All five modules are on-demand video with notebooks you run yourself. Start whenever you want, go at your own speed, and keep lifetime access to everything. Questions go to the private community.
Give Your Software Eyes
Go from “VLMs sound interesting” to a working vision pipeline in 5 hours, at your own pace.
Join the WaitlistCourse launches June 2026. Join the waitlist to get notified.