Self-Paced Course · June 2026

Computer Visionwith VLMs

An introduction to vision-language models for builders. Understand how modern AI sees images — and ship your first vision pipeline in 5 hours.

Self-Paced5 Hours TotalNo ML Background Needed

How a VLM answers

Image → patches

The image is split into a grid of tokens

Vision encoder

Patches become embeddings the LLM can read

“The invoice total is $1,240.”

Language model reasons over both

What You'll Walk Away With

Short and dense by design. Built for engineers and technical PMs who want to add vision to their products without a research detour.

A working mental model

Understand what happens between an image going in and an answer coming out — so failures stop being mysterious.

Practical prompting skills

Extraction, grounding, comparison, and multi-image workflows you can apply the same day at work.

One real pipeline

Leave with a document-understanding pipeline you built yourself, ready to adapt to your own data.

The 5-Hour Curriculum

Five focused modules you complete at your own pace. Every module ends with something you built.

01

From Pixels to Meaning

45 min

Why classic computer vision hit a wall, and how vision-language models changed what "seeing" means for software. CLIP, contrastive learning, and the core intuition you need.

02

How a VLM Actually Sees

60 min

Image patches, vision encoders, and projection into the language space. No math beyond what you need — just enough to predict when a model will fail.

03

Prompting for Vision Tasks

75 min

Hands-on: visual question answering, grounding and bounding boxes, structured extraction from images, and multi-image reasoning. Working with GPT-4o, Claude, and Qwen-VL.

04

Real Applications

60 min

OCR-free document parsing, UI and screenshot understanding, product catalog tagging, and video frame analysis. Build one working pipeline end-to-end.

05

Limits, Evals & Model Choice

60 min

Hallucinated details, counting failures, spatial reasoning gaps. How to eval a VLM for your use case and pick between hosted APIs and open weights.

Join the Waitlist

Be the first to know when the course launches in June 2026.

Start whenever you want — learn at your own pace

Course launches
June 2026
5 On-Demand Video Modules (5h total)
Lifetime Access to All Materials
Working Vision Pipeline Notebook
Certificate of Completion
Private Discord Community
Q&A in the Community

Questions? Contact us at hello@theagentcamp.com

Frequently Asked Questions

Do I need machine learning experience?

No. We explain the internals at the level of intuition, not equations. If you can write Python and call an API, you have everything you need.

Which models will we use?

We work hands-on with hosted APIs (GPT-4o, Claude) and one open-weights model (Qwen-VL) so you see both sides of the trade-off.

Is 5 hours really enough?

For an introduction, yes. The goal is a solid mental model and one working pipeline — not covering every architecture. You leave knowing exactly what to learn next.

How does self-paced work?

All five modules are on-demand video with notebooks you run yourself. Start whenever you want, go at your own speed, and keep lifetime access to everything. Questions go to the private community.

Give Your Software Eyes

Go from “VLMs sound interesting” to a working vision pipeline in 5 hours, at your own pace.

Join the Waitlist

Course launches June 2026. Join the waitlist to get notified.