This course covers vision-language models and multi-agent coordination: how agents interpret images alongside text and divide work between them. Together they take an agent beyond single inputs.
You explore how vision-language models such as CLIP, BLIP, and LLaVA align images with text, and use those embeddings to build a multimodal search system. You then construct a multimodal RAG pipeline that answers questions about charts and tables inside PDF documents, where text-only retrieval fails. The course closes with orchestration: planner, executor, and critic patterns that break complex tasks into steps, reliable tool schemas with error handling, and collaborative multi-agent systems where specialized agents pass context between each other. By the end of this course, you will be able to: - Explain how vision-language models align image and text representations. - Build a multimodal search system using shared embedding spaces. - Construct a multimodal RAG pipeline over documents containing charts. - Apply planner, executor, and critic patterns to decompose complex tasks. - Implement tool calling with reliable schemas and error handling. - Coordinate multiple agents through defined roles, handoffs, and shared context. Intended for learners who have completed Multimodal AI and Agent Fundamentals. Enroll now to give your agents sight, retrieval, and the ability to work as a team.













