Antoine Amend is a data practitioner passionate about distributed computing and advanced analytics. Graduated with a master degree in computational astrophysics, author of “Mastering Spark for data science”, Antoine has been pushing both engineering and science disciplines side by side to extract commercial value from large datasets. More recently, Antoine served as director of data science at Barclays UK, leading their AI practice and driving Barclays through their data and analytics transformation. With his expertise in enterprise architecture and commercial experience delivering data science to production in a highly regulated environment, Antoine joined Databricks as the technical director for financial services, helping our customers redefine the future of Banking.
We demonstrate how to use Spark Streaming to build a global News Scanner that scrapes news in near real time, and uses sophisticated text analysis, SimHash, Random Indexing and Streaming K-Means to produce a geopolitical monitoring tool that allows users to track major world events as they unfold. We highlight advanced spark techniques for scaling, including: using Apache NIFI to deliver data to Spark Streaming, using the Goose library with Spark to build web scrapers, how to de-duplicate streamed documents at scale using advanced techniques like SimHash, Random Indexing, and Streaming K-Means in order to detect, track and visualise "global media conversations” as they mutate over time. Session hashtag: #EUstr9