Skip to main content

Databricks Labs

Databricks Labs are projects created by the field team to help customers get their use cases into production faster!

Node

dqx

DQX

Simplified Data Quality checking at Scale for PySpark Workloads on streaming and standard DataFrames.

GitHub Sources →

Documentation →

Kasal

sdp-meta

A metadata-driven framework for Lakeflow Spark Declarative Pipelines. Define Bronze and Silver pipelines in a JSON or YAML onboarding file — a single generic pipeline reads the resulting spec at runtime and builds the full processing graph, so a single data engineer can manage thousands of tables without writing new pipeline code. Available via Bundles, CLI, a browser-based App, and an MCP server for AI-assisted setup.

Github Sources →

Documentation →

Interview with Author →

Lakebridge

GeoBrix

GeoBrix delivers high-performance spatial processing that complements Databricks' native spatial capabilities — GEOMETRY/GEOGRAPHY types, ST_* functions, and H3. It extends the Lakehouse with advanced raster processing, vector and raster tile generation, enhanced geometry tools, and additional Discrete Global Grids such as Quadbin and the British National Grid. Its lightweight tier runs friction-free on Databricks Serverless and Lakeflow pipelines, and a suite of specialized readers and writers adds format-specific handling for common geospatial file types.

Github Sources →

Documentation →

Blog →

Other Projects

Kasal

Kasal is an interactive, low-code way to build and deploy AI Agents on the Databricks platform.

Github Sources →
Documentation →

Lakebridge

Lakebridge is Databricks’ migration platform, designed to provide enterprises with a comprehensive, end-to-end solution for modernizing legacy data warehouses and ETL systems. Lakebridge supports a wide range of source platforms —including Teradata, Oracle, Snowflake, SQL Server, DataStage, and more— and automates every stage of the migration process, from discovery and assessment to code conversion, data movement, and validation, ensuring a fast, low-risk transition for organizations seeking to unlock innovation and efficiency in their data estate.

GitHub Sources →
Documentation →
Blog →

Impulse

Impulse is a Python-based analytics library designed for processing large-scale time-series measurement data. Built on Apache Spark and Delta Lake, it enables distributed processing of petabyte-scale sensor data from automotive testing, industrial IoT, and other measurement-intensive domains.

Github Sources →
Documentation →

OntoBricks

OntoBricks is a web application that transforms Databricks tables into a materialized graph viewer. It lets you design ontologies (OWL), map them to Unity Catalog tables via R2RML, materialize triples into a Delta-backed triple store and a Lakebase Postgres graph engine, reason over the graph (OWL 2 RL, SWRL, SHACL), and query it through an auto-generated GraphQL API. The entire pipeline—from metadata import to a queryable graph viewer—can run in four clicks using LLM-powered automation.

Github Sources →
Documentation →

coda

Run Claude Code, Codex, Gemini CLI, Hermes Agent, and OpenCode in your browser—zero setup, wired to your Databricks workspace.

Github Sources →
Documentation →

VibeScaler

Collaborate with your team to define what good agent behavior looks like, then turn that judgment into an automated grader that runs at scale.

GitHub Sources →
Documentation →

Lakemeter

Estimate Databricks workload costs in minutes—with transparent assumptions you can review, share, and export.

Github Sources →
Documentation →

Ontos

Ontos provides enterprise teams with the tools to organize, govern, and deliver high-quality data products following Data Mesh principles and industry standards like ODCS (Open Data Contract Standard) and ODPS (Open Data Product Specification).

GitHub Sources →
Documentation →
Marketplace →

Databricks MCP

A collection of MCP servers to help AI agents fetch enterprise data from Databricks and automate common developer actions on Databricks.

Github Sources →

Conversational Agent App

Application featuring a chat interface powered by Databricks Genie Conversation APIs, built specifically to run as a Databricks App.

Github Sources →

Knowledge Assistant Chatbot Application

Example Databricks Knowledge Assistant chatbot application.

Github Sources →

Feature Registry Application

The app provides a user-friendly interface for exploring existing features in Unity Catalog. Additionally, users can generate code for creating feature specs and training sets to train machine learning models and deploy features as Feature Serving Endpoints.

Github Sources →

Smolder

Smolder provides an Apache Spark™ SQL data source for loading EHR data from HL7v2 message formats. Additionally, Smolder provides helper functions that can be used on a Spark SQL DataFrame to parse HL7 message text, and to extract segments, fields, and subfields from a message.

Github Sources →
Learn more →

Data Generator

Generate relevant data quickly for your projects. The Databricks data generator can be used to generate large simulated/synthetic data sets for test, POCs, and other uses

Github Sources →
Learn more →

DeltaOMS

Centralized Delta transaction log collection for metadata and operational metrics analysis on your Lakehouse.

Github Sources →
Learn more →

Splunk Integration

Add-on for Splunk, an app that allows Splunk Enterprise and Splunk Cloud users to run queries and execute actions, such as running notebooks and jobs, in Databricks.

Github Sources →
Learn more →

DiscoverX

DiscoverX automates administration tasks that require inspecting or applying operations to a large number of Lakehouse assets.

Github Sources →

brickster

{brickster} is the R toolkit for Databricks, it includes:

  • Wrappers for Databricks API's (e.g. db_cluster_list, db_volume_read)
  • Browser workspace assets via RStudio Connections Pane (open_workspace())
  • Exposes the databricks-sql-connector via {reticulate} (docs)
  • Interactive Databricks REPL

Github Sources →
Documentation →
Blog →

Tempo

The purpose of this project is to provide an API for manipulating time series on top of Apache Spark™. Functionality includes featurization using lagged time values, rolling statistics (mean, avg, sum, count, etc.), AS OF joins, and downsampling and interpolation. This has been tested on TB-scale of historical data.

GitHub Sources →
Documentation →
Webinar →

PyLint Plugin

This plugin extends PyLint with checks for common mistakes and issues in Python code specifically in Databricks Environment.

Github Sources →
Documentation →

PyTester

PyTester is a powerful way to manage test setup and teardown in Python. This library provides a set of fixtures to help you write integration tests for Databricks.

Github Sources →
Documentation →

Delta Sharing Java Connector

The Java connector follows the Delta Sharing protocol to read shared tables from a Delta Sharing Server. To further reduce and limit egress costs on the Data Provider side, we implemented a persistent cache to reduce and limit the egress costs on the Data Provider side by removing any unnecessary reads.

GitHub Sources →
Documentation →

UCX

UCX is a toolkit for enabling Unity Catalog (UC) in your Databricks workspace. UCX provides commands and workflows for migrate tables and views to UC. UCX allows to rewrite dashboards, jobs and notebooks to use the migrated data assets in UC. And there are many more features.

GitHub Sources →
Documentation →
Blog →

migrate-ip-acls

Recreate a Databricks workspace's existing IP access list as a context-based ingress (CBI) policy via a single, focused CLI.

GitHub Sources →

Firefly

A Next.js application that provides a customized frontend for Databricks with multiple authentication strategies and embedded Databricks apps.

Github Sources →
Documentation →

Please note that all projects in the https://github.com/databrickslabs account are provided for your exploration only, and are not formally supported by Databricks with service level agreements (SLAs). They are provided AS IS and we do not make any guarantees of any kind.  Any issues discovered through the use of these projects can be filed as GitHub Issues on the Repo. They will be reviewed as time permits, but there are no formal SLAs for GitHub support.