IBM

Unstructured Data Engineering for AI

Ce cours n'est pas disponible en Français (France)

Nous sommes actuellement en train de le traduire dans plus de langues.
IBM

Unstructured Data Engineering for AI

Antonio Cangiano
Ruslan Podgaets

Instructeurs : Antonio Cangiano

Inclus avec Coursera PlusEn savoir plus

Demander à Coursera

Obtenez un aperçu d'un sujet et apprenez les principes fondamentaux.
niveau Intermédiaire

Expérience recommandée

2 semaines à compléter
à 10 heures par semaine
Planning flexible
Apprenez à votre propre rythme
Obtenez un aperçu d'un sujet et apprenez les principes fondamentaux.
niveau Intermédiaire

Expérience recommandée

2 semaines à compléter
à 10 heures par semaine
Planning flexible
Apprenez à votre propre rythme

Ce que vous apprendrez

  • 1.Build ingestion and extraction workflows for documents and multimodal AI corpora.

  • 2. Apply OCR-aware processing, normalization, and PII-safe corpus preparation.

  • 3. Design chunking and metadata enrichment for retrieval, training, and citation use cases.

  • 4. Apply safety, licensing, bias, and source-quality gates to unstructured AI data assets.

Compétences que vous acquerrez

  • Catégorie : Data Collection
  • Catégorie : Data Processing
  • Catégorie : Responsible AI
  • Catégorie : Retrieval-Augmented Generation
  • Catégorie : Document Management
  • Catégorie : Release Management
  • Catégorie : Metadata Management
  • Catégorie : Data Cleansing
  • Catégorie : Information Architecture
  • Catégorie : Taxonomy
  • Catégorie : Data Quality
  • Catégorie : Data Capture
  • Catégorie : Data Governance
  • Catégorie : Personally Identifiable Information
  • Catégorie : Data Engineering
  • Catégorie : Text Mining
  • Catégorie : Unstructured Data
  • Catégorie : Data Architecture

Outils que vous découvrirez

  • Catégorie : Data Lakes
  • Catégorie : AI Workflows

Détails à connaître

Certificat partageable

Ajouter à votre profil LinkedIn

Récemment mis à jour !

juillet 2026

Évaluations

30 devoirs

Enseigné en Anglais

Découvrez comment les employés des entreprises prestigieuses maîtrisent des compétences recherchées

 logos de Petrobras, TATA, Danone, Capgemini, P&G et L'Oreal

Élaborez votre expertise en Machine Learning

Ce cours fait partie de la Certificat Professionnel IBM AI-Native Data Engineering
Lorsque vous vous inscrivez à ce cours, vous êtes également inscrit(e) à ce Certificat Professionnel.
  • Apprenez de nouveaux concepts auprès d'experts du secteur
  • Acquérez une compréhension de base d'un sujet ou d'un outil
  • Développez des compétences professionnelles avec des projets pratiques
  • Obtenez un certificat professionnel partageable auprès de IBM

Il y a 9 modules dans ce cours

This welcome module introduces Course 5 and explains why unstructured data engineering matters for AI-native data work. Learners will orient themselves to the course’s professional value, expected preparation, and where to find the full course roadmap before beginning the technical modules.

Inclus

1 vidéo2 plugins

Learn how to design a governed object storage foundation for unstructured AI datasets, including storage zones, naming conventions, manifests, and file quality controls. By the end of the module, you will be able to prepare an ingestion-ready corpus that supports downstream extraction, governance, lineage, and auditability.

Inclus

4 vidéos4 devoirs2 éléments d'application4 plugins

Learn how to extract text and preserve useful structure from PDFs, Word files, HTML, Markdown, and scanned documents while deciding when OCR is needed. You will build an extraction workflow, add quality checks and routing logic, and package audit-ready outputs for downstream cleaning, chunking, and governed AI use.

Inclus

4 vidéos5 devoirs2 éléments d'application4 plugins

Learn how to turn extracted unstructured text into cleaner, safer, and more reliable corpus content for downstream AI workflows. You will inspect and normalize noisy text, remove boilerplate without losing important structure, and apply sensitive data detection, redaction, and audit practices that support governed corpus release.

Inclus

4 vidéos5 devoirs2 éléments d'application4 plugins

This module teaches learners how to turn cleaned unstructured content into AI-ready chunks enriched with metadata for traceability, filtering, citation readiness, governance, and auditability. Learners compare chunking strategies, define chunk schemas, enrich records with source and governance metadata, and validate chunk quality before downstream embedding, retrieval, or training workflows.

Inclus

4 vidéos4 devoirs2 éléments d'application4 plugins

This module teaches learners to design governed annotation and labeling workflows for unstructured AI corpora, from taxonomy creation through human review, agreement checks, AI-assisted labeling boundaries, versioning, and governance metadata. By the end, learners will be able to produce auditable labeling artifacts that improve corpus quality, reproducibility, and downstream AI readiness.

Inclus

4 vidéos5 devoirs2 éléments d'application4 plugins

This module brings together artifacts from earlier modules to evaluate whether an unstructured corpus is truly ready for AI use. You will define and apply safety, quality, licensing, and source trust gates, interpret validation results, and package a governed corpus release with clear evidence and a professional final report.

Inclus

4 vidéos5 devoirs4 plugins

This short wrap-up module closes Course 5 by helping learners reflect on the value of governed unstructured data engineering for AI and recognize the progress they have made. It also previews how Course 6 builds on these themes without introducing new technical content.

Inclus

1 vidéo1 plugin

This Final Exam assesses your ability to apply the full unstructured data engineering workflow for AI, from ingestion and extraction to normalization, chunking, labeling, and governed release decisions. You will demonstrate both core knowledge and practical judgment about traceable, reproducible, and risk aware corpus pipeline design.

Inclus

2 devoirs1 plugin

Obtenez un certificat professionnel

Ajoutez ce titre à votre profil LinkedIn, à votre curriculum vitae ou à votre CV. Partagez-le sur les médias sociaux et dans votre évaluation des performances.

Instructeurs

Antonio Cangiano
IBM
15 Cours751 278 apprenants
Ruslan Podgaets
IBM
5 Cours73 apprenants

Offert par

IBM

En savoir plus sur Machine Learning

Pour quelles raisons les étudiants sur Coursera nous choisissent-ils pour leur carrière ?

Felipe M.

Étudiant(e) depuis 2018
’Pouvoir suivre des cours à mon rythme à été une expérience extraordinaire. Je peux apprendre chaque fois que mon emploi du temps me le permet et en fonction de mon humeur.’

Jennifer J.

Étudiant(e) depuis 2020
’J'ai directement appliqué les concepts et les compétences que j'ai appris de mes cours à un nouveau projet passionnant au travail.’

Larry W.

Étudiant(e) depuis 2021
’Lorsque j'ai besoin de cours sur des sujets que mon université ne propose pas, Coursera est l'un des meilleurs endroits où se rendre.’

Chaitanya A.

’Apprendre, ce n'est pas seulement s'améliorer dans son travail : c'est bien plus que cela. Coursera me permet d'apprendre sans limites.’

Foire Aux Questions