# From Raw Documents to AI-Ready Data — Leo Platzer & Jeff Koss

**URL:** <https://www.artofsm.art/t/from-raw-documents-to-ai-ready-data-leo-platzer-jeff-koss/25095>\
**Category:** Content Creators\
**Tags:** chatbot, ai-engineer, automation\
**Created:** [5 October 2026 22:30 UTC](https://www.artofsm.art/t/from-raw-documents-to-ai-ready-data-leo-platzer-jeff-koss/25095 "2026-10-05T22:30:12Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![artesia](https://www.artofsm.art/user_avatar/www.artofsm.art/artesia/32/36_2.png) [@artesia](https://www.artofsm.art/u/artesia)\
**Post date:** [5 October 2026 22:30 UTC](https://www.artofsm.art/t/from-raw-documents-to-ai-ready-data-leo-platzer-jeff-koss/25095/1 "2026-10-05T22:30:12Z")

</div>

[![](https://www.artofsm.art/uploads/default/original/3X/6/c/6c16e625e463ecb09e2eee11f8614632b74eec5f.jpeg "From Raw Documents to AI-Ready Data — Leo Platzer & Jeff Koss") ](https://www.youtube.com/watch?v=wzWNYDY7toc)

Jeff Koss and Leo Platzer from Deasy Labs present a solution for transforming vast amounts of unstructured documents into AI-ready data, enhancing chatbot accuracy and operational efficiency by addressing challenges like data quality, duplicates, and sensitivity management. Their platform automates tagging, metadata generation, and ongoing data curation, significantly improving AI performance for complex legal and procurement queries while reducing compliance risks and accelerating data preparation timelines.

---

<div class="post-metadata">

**Author:** ![artesia](https://www.artofsm.art/user_avatar/www.artofsm.art/artesia/32/36_2.png) [@artesia](https://www.artofsm.art/u/artesia)\
**Post date:** [5 October 2026 22:51 UTC](https://www.artofsm.art/t/from-raw-documents-to-ai-ready-data-leo-platzer-jeff-koss/25095/3 "2026-10-05T22:51:02Z")

</div>

The presentation, led by Jeff Koss and Leo Platzer from Deasy Labs (now part of Collibra), focuses on transforming raw, unstructured documents into AI-ready data to enhance chatbot performance and operational efficiency. They introduce a manufacturing company, North River Manufacturing, which faces challenges managing millions of documents across multiple countries, aiming to develop an AI chatbot to assist legal and procurement teams with real-time queries about contracts and agreements. The team highlights key stakeholders involved, including data scientists, AI engineers, and legal operations directors, each with specific needs around data accuracy, relevance, and confidentiality.

The core challenge discussed is scaling from a small pilot project with 40 files to handling over 80,000 documents, revealing issues such as difficulty locating relevant files, managing sensitive information, and ensuring data quality. Jeff emphasizes that traditional data quality dimensions like completeness and accuracy must be adapted for unstructured data, focusing on duplicates, conflicting information, and data freshness. These factors critically impact chatbot accuracy and reliability, as outdated or redundant information can degrade AI responses and increase risks like data leaks or compliance failures.

The demo showcases Deasy Labs’ platform capabilities, including automated tagging using AI, taxonomy creation, metadata generation, and sensitivity scanning for PII or confidential data. Users can filter and classify documents by relevance and legal category, with AI providing evidence for tagging decisions and allowing human feedback to improve accuracy. The system supports ongoing data quality management through dashboards that identify duplicates and conflicting facts, and it enables scheduled workflows to keep the AI data product current by automatically updating with new or removed files.

Leo expands on the importance of managing duplicates and data freshness, explaining how outdated or conflicting information can flood the AI’s knowledge base, reducing the effectiveness of chatbot responses. He shares evaluation results showing that cleaning and curating unstructured data can nearly double recall and improve accuracy by 10-15% in complex use cases. Leo also demonstrates how their SDK can create contextual metadata repositories for SharePoint folders, allowing AI systems to quickly understand document contents and relevance without scanning every file, thereby optimizing performance and reducing computational costs.

In conclusion, Jeff invites attendees to engage with their team for live demos and personalized proof-of-concept sessions, emphasizing the value of their technology in accelerating data preparation from months to days, improving AI chatbot accuracy, and mitigating compliance risks. They offer ongoing support to tailor taxonomies and metadata generation to specific business needs, helping organizations unlock the potential of their unstructured data for AI applications. The session closes with an invitation to visit their booth for further discussions and giveaways.
