IBM

Unstructured Data Engineering for AI

kurs ist nicht verfügbar in Deutsch (Deutschland)

Wir übersetzen es in weitere Sprachen.
IBM

Unstructured Data Engineering for AI

Antonio Cangiano
Ruslan Podgaets

Dozenten: Antonio Cangiano

Bei Coursera PlusMehr erfahren enthalten

Verschaffen Sie sich einen Einblick in ein Thema und lernen Sie die Grundlagen.
Stufe Mittel

Empfohlene Erfahrung

2 Wochen zu vervollständigen
unter 10 Stunden pro Woche
Flexibler Zeitplan
In Ihrem eigenen Lerntempo lernen
Verschaffen Sie sich einen Einblick in ein Thema und lernen Sie die Grundlagen.
Stufe Mittel

Empfohlene Erfahrung

2 Wochen zu vervollständigen
unter 10 Stunden pro Woche
Flexibler Zeitplan
In Ihrem eigenen Lerntempo lernen

Was Sie lernen werden

  • 1.Build ingestion and extraction workflows for documents and multimodal AI corpora.

  • 2. Apply OCR-aware processing, normalization, and PII-safe corpus preparation.

  • 3. Design chunking and metadata enrichment for retrieval, training, and citation use cases.

  • 4. Apply safety, licensing, bias, and source-quality gates to unstructured AI data assets.

Kompetenzen, die Sie erwerben

  • Kategorie: Release Management
  • Kategorie: Document Management
  • Kategorie: Metadata Management
  • Kategorie: Information Architecture
  • Kategorie: Data Cleansing
  • Kategorie: Retrieval-Augmented Generation
  • Kategorie: Responsible AI
  • Kategorie: Data Engineering
  • Kategorie: Data Collection
  • Kategorie: Data Processing
  • Kategorie: Data Capture
  • Kategorie: Data Governance
  • Kategorie: Personally Identifiable Information
  • Kategorie: Data Quality
  • Kategorie: Text Mining
  • Kategorie: Unstructured Data
  • Kategorie: Taxonomy
  • Kategorie: Data Architecture

Werkzeuge, die Sie lernen werden

  • Kategorie: Data Lakes
  • Kategorie: AI Workflows

Wichtige Details

Zertifikat zur Vorlage

Zu Ihrem LinkedIn-Profil hinzufügen

Kürzlich aktualisiert!

Juli 2026

Bewertungen

30 Aufgaben

Unterrichtet in Englisch

Erfahren Sie, wie Mitarbeiter führender Unternehmen gefragte Kompetenzen erwerben.

 Logos von Petrobras, TATA, Danone, Capgemini, P&G und L'Oreal

Erweitern Sie Ihr Fachwissen im Bereich Machine Learning

Dieser Kurs ist Teil der Spezialisierung IBM AI-Native Data Engineering (berufsbezogenes Zertifikat)
Wenn Sie sich für diesen Kurs anmelden, werden Sie auch für dieses berufsbezogene Zertifikat angemeldet.
  • Lernen Sie neue Konzepte von Branchenexperten
  • Gewinnen Sie ein Grundverständnis bestimmter Themen oder Tools
  • Erwerben Sie berufsrelevante Kompetenzen durch praktische Projekte
  • Erwerben Sie ein Berufszertifikat von IBM zur Vorlage

In diesem Kurs gibt es 9 Module

This welcome module introduces Course 5 and explains why unstructured data engineering matters for AI-native data work. Learners will orient themselves to the course’s professional value, expected preparation, and where to find the full course roadmap before beginning the technical modules.

Das ist alles enthalten

1 Video2 Plug-ins

Learn how to design a governed object storage foundation for unstructured AI datasets, including storage zones, naming conventions, manifests, and file quality controls. By the end of the module, you will be able to prepare an ingestion-ready corpus that supports downstream extraction, governance, lineage, and auditability.

Das ist alles enthalten

4 Videos4 Aufgaben2 App-Elemente4 Plug-ins

Learn how to extract text and preserve useful structure from PDFs, Word files, HTML, Markdown, and scanned documents while deciding when OCR is needed. You will build an extraction workflow, add quality checks and routing logic, and package audit-ready outputs for downstream cleaning, chunking, and governed AI use.

Das ist alles enthalten

4 Videos5 Aufgaben2 App-Elemente4 Plug-ins

Learn how to turn extracted unstructured text into cleaner, safer, and more reliable corpus content for downstream AI workflows. You will inspect and normalize noisy text, remove boilerplate without losing important structure, and apply sensitive data detection, redaction, and audit practices that support governed corpus release.

Das ist alles enthalten

4 Videos5 Aufgaben2 App-Elemente4 Plug-ins

This module teaches learners how to turn cleaned unstructured content into AI-ready chunks enriched with metadata for traceability, filtering, citation readiness, governance, and auditability. Learners compare chunking strategies, define chunk schemas, enrich records with source and governance metadata, and validate chunk quality before downstream embedding, retrieval, or training workflows.

Das ist alles enthalten

4 Videos4 Aufgaben2 App-Elemente4 Plug-ins

This module teaches learners to design governed annotation and labeling workflows for unstructured AI corpora, from taxonomy creation through human review, agreement checks, AI-assisted labeling boundaries, versioning, and governance metadata. By the end, learners will be able to produce auditable labeling artifacts that improve corpus quality, reproducibility, and downstream AI readiness.

Das ist alles enthalten

4 Videos5 Aufgaben2 App-Elemente4 Plug-ins

This module brings together artifacts from earlier modules to evaluate whether an unstructured corpus is truly ready for AI use. You will define and apply safety, quality, licensing, and source trust gates, interpret validation results, and package a governed corpus release with clear evidence and a professional final report.

Das ist alles enthalten

4 Videos5 Aufgaben4 Plug-ins

This short wrap-up module closes Course 5 by helping learners reflect on the value of governed unstructured data engineering for AI and recognize the progress they have made. It also previews how Course 6 builds on these themes without introducing new technical content.

Das ist alles enthalten

1 Video1 Plug-in

This Final Exam assesses your ability to apply the full unstructured data engineering workflow for AI, from ingestion and extraction to normalization, chunking, labeling, and governed release decisions. You will demonstrate both core knowledge and practical judgment about traceable, reproducible, and risk aware corpus pipeline design.

Das ist alles enthalten

2 Aufgaben1 Plug-in

Erwerben Sie ein Karrierezertifikat.

Fügen Sie dieses Zeugnis Ihrem LinkedIn-Profil, Lebenslauf oder CV hinzu. Teilen Sie sie in Social Media und in Ihrer Leistungsbeurteilung.

Dozenten

Antonio Cangiano
IBM
15 Kurse751.278 Lernende
Ruslan Podgaets
IBM
5 Kurse73 Lernende

von

IBM

Mehr von Machine Learning entdecken

Warum entscheiden sich Menschen für Coursera für ihre Karriere?

Felipe M.

Lernender seit 2018
„Es ist eine großartige Erfahrung, in meinem eigenen Tempo zu lernen. Ich kann lernen, wenn ich Zeit und Nerven dazu habe.“

Jennifer J.

Lernender seit 2020
„Bei einem spannenden neuen Projekt konnte ich die neuen Kenntnisse und Kompetenzen aus den Kursen direkt bei der Arbeit anwenden.“

Larry W.

Lernender seit 2021
„Wenn mir Kurse zu Themen fehlen, die meine Universität nicht anbietet, ist Coursera mit die beste Alternative.“

Chaitanya A.

„Man lernt nicht nur, um bei der Arbeit besser zu werden. Es geht noch um viel mehr. Bei Coursera kann ich ohne Grenzen lernen.“

Häufig gestellte Fragen