Anonymization of sensitive information in financial documents (sps25) episode artwork

EPISODE · Oct 17, 2025 · 31 MIN

Anonymization of sensitive information in financial documents (sps25)

from Chaos Computer Club - recent events feed (high quality) · host Piotr Gryko

Data is the fossil fuel of the machine learning world, essential for developing high quality models but in limited supply. Yet institutions handling sensitive documents — such as financial, medical, or legal records often cannot fully leverage their own data due to stringent privacy, compliance, and security requirements, making training high quality models difficult. A promising solution is to replace the personally identifiable information (PII) with realistic synthetic stand-ins, whilst leaving the rest of the document in tact. In this talk, we will discuss the use of open source tools and models that can be self hosted to anonymize documents. We will go over the various approaches for Named Entity Recognition (NER) to identify sensitive entities and the use of diffusion models to inpaint anonymized content. about this event: https://talks.python-summit.ch/sps25/talk/EWMBRM/

Episode metadata supplied by the publisher feed · Published Oct 17, 2025

Embed this episode

NOW PLAYING

Anonymization of sensitive information in financial documents (sps25)

0:00 31:17

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Chaos Computer Club - recent events feed (high quality)?

This episode is 31 minutes long.

When was this Chaos Computer Club - recent events feed (high quality) episode published?

This episode was published on October 17, 2025.

Can I download this Chaos Computer Club - recent events feed (high quality) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!