Machine Learning at Scale

In this course, you will gain theoretical and practical knowledge of Apache Spark’s architecture and its application to machine learning workloads within Databricks. You will learn when to use Spark for data preparation, model training, and deployment, while also gaining hands-on experience with Spark ML and pandas APIs on Spark. This course will introduce you to advanced concepts like hyperparameter tuning and scaling Optuna with Spark. This course will use features and concepts introduced in the associate course such as MLflow and Unity Catalog for comprehensive model packaging and governance.

Note: This course is the first in the series of Advanced Machine Learning.

Skill Level

Professional

Duration

Prerequisites

The content was developed for participants with these skills/knowledge/abilities:

• Familiarity with the Databricks Data Intelligence Platform and basic workspace operations (create clusters, run code in notebooks, use basic notebook operations, import repos from git).

• Intermediate programming experience with Python, including data manipulation libraries (pandas, numpy) and machine learning frameworks (scikit-learn).

• Basic knowledge of Apache Spark and PySpark fundamentals, including DataFrames, transformations, and actions for distributed data processing.

• Understanding of machine learning concepts, including model training, evaluation, hyperparameter tuning, and deployment workflows.

• Intermediate experience with Delta Lake operations (create tables, perform updates, optimize files, time travel functionality).

• Basic familiarity with MLflow for experiment tracking, model logging, and model registry operations.

• Understanding of distributed computing concepts (cluster architecture, parallelization, scalability considerations).

• Basic knowledge of SQL for data querying and manipulation within Spark environments.

Self-Paced

Custom-fit learning paths for data, analytics, and AI roles and career paths through on-demand videos

Customer registration Partner registration

See all our registration options

Registration options

Databricks has a delivery method for wherever you are on your learning journey

Self-Paced

Custom-fit learning paths for data, analytics, and AI roles and career paths through on-demand videos

Instructor-Led

Public and private courses taught by expert instructors across half-day to two-day courses

Blended Learning

Self-paced and weekly instructor-led sessions for every style of learner to optimize course completion and knowledge retention. Go to Subscriptions Catalog tab to purchase

Purchase now

Skills@Scale

Comprehensive training offering for large scale customers that includes learning elements for every style of learning. Inquire with your account executive for details

Upcoming Public Classes

Data Analyst

SQL Analytics on Databricks

In this course, you'll learn how to effectively use Databricks for data analytics, with a specific focus on Databricks SQL. As a Databricks Data Analyst, your responsibilities will include finding relevant data, analyzing it for potential applications, and transforming it into formats that provide valuable business insights. You will also understand your role in managing data objects and how to manipulate them within the Databricks Data Intelligence Platform, using tools such as Notebooks, the SQL Editor, and Databricks SQL. Additionally, you will learn about the importance of Unity Catalog in managing data assets and the overall platform. Finally, the course will provide an overview of how Databricks facilitates performance optimization and teach you how to access Query Insights to understand the processes occurring behind the scenes when executing SQL analytics on Databricks.

Note: Databricks Academy is transitioning from video lectures to a more streamlined PDF format with slides and notes for all self-paced courses. Please note that demo videos will still be available in their original format. We would love to hear your thoughts on this change, so please share your feedback through the course survey at the end. Thank you for being a part of our learning community!

Languages Available: English | 日本語 | Português BR | 한국어

Paid & Subscription

Lab

Associate

Apache Spark Developer

Introduction to Apache Spark™

This course offers essential knowledge of Apache Spark™, with a focus on its distributed architecture and practical applications for large-scale data processing. Participants will explore programming frameworks, learn the Spark DataFrame API, and develop skills for reading, writing, and transforming data using Python-based Spark workflows.

Languages Available: English | 한국어

Paid & Subscription

Lab

Associate

Agent Evaluation on Databricks

This course teaches students how to systematically evaluate AI agents using MLflow's evaluation framework, addressing the unique challenges of non-deterministic AI systems that traditional software testing cannot handle. Students learn to implement various evaluation approaches including built-in judges for common criteria like correctness and safety, guideline judges for business-specific requirements, and custom judges for specialized needs. The course covers both offline evaluation using curated datasets and online production monitoring, with hands-on experience using MLflow's tracing capabilities to understand agent execution patterns and collect human feedback from different stakeholder types. Through practical demonstrations and labs, students develop skills in creating evaluation workflows that drive continuous quality improvements throughout the AI agent development lifecycle.

Paid & Subscription

Lab

Associate