You can finally BLOCK GENIE usage when a budget limit is reached

You can finally BLOCK GENIE usage when a budget limit is reached

I’ve been testing the new Genie budget controls, and this is probably the most important addition since Genie Code moved to pay-as-you-go. Previously, budgets were mostly useful for monitoring and alerts. Now you can actually block Genie usage when a spending threshold is reached.

A few things I found useful while testing:

  • shared budget + per-user thresholds
  • different overrides for users/groups
  • email notification or BLOCK_USAGE
  • actual Genie consumption can be checked in system.billing.usage
  • free usage is visible under GENIE_FREE_USAGE

For admins, this makes Genie much easier to roll out to a large number of users without leaving AI spending effectively open-ended.

I put the setup, SQL and my tests here: https://medium.com/databrickscommunity/databricks-genie-cost-control-how-to-set-budgets-and-block-usage-a13014c1f9ba

u/Significant-Guest-14 — 3 days ago

Databricks Introduces Spaces for Organizing Your Development Context

Databricks has Spaces again 😄. Not Genie Spaces, those were recently renamed to Genie Agents. This is a completely different feature.

Spaces let you save your development context in Databricks:

  • Project folder
  • Open tabs
  • Scoped search
  • Working context

So if you're working on multiple projects, you can have a Space for each one and switch between them without reopening notebooks and navigating back to the right folders. Spaces also work with Git folders, Lakeflow Pipelines, and Declarative Automation Bundles. Nothing revolutionary, but it looks like a useful quality-of-life improvement if you spend a lot of time in the Databricks workspace.

Genie Spaces → Genie Agents
New feature → Spaces (for Notebooks)

oficcial docs: https://docs.databricks.com/aws/en/notebooks/spaces

u/Significant-Guest-14 — 8 days ago

Follow-up: Centralized Medallion vs Data Mesh vs Hybrid Data Mesh

Following up on yesterday's Data Mesh discussion on Reddit.

The comments pointed out a fair issue with my original architecture: what I showed wasn't a classical Data Mesh. So I updated the article, renamed the approach to Hybrid Data Mesh, and simplified the diagram to compare three models:

  • Centralized Medallion: Central Data Team owns Bronze → Silver → Gold.
  • Data Mesh: Domain Teams own their pipelines and Data Products end-to-end.
  • Hybrid Data Mesh: Central Data Team owns the common platform and Bronze/Silver, while Domain Teams own Gold and business logic.

For me, the interesting question isn't which one is the "correct" Data Mesh. It's: Where should the Central Data Team's ownership end and Domain ownership begin?

Updated article: https://medium.com/databrickscommunity/databricks-data-mesh-best-practices-a-practical-implementation-guide-b54309bc5f3e

u/Significant-Guest-14 — 9 days ago

I couldn't find a practical Data Mesh guide for Databricks, so I wrote one

I've been looking for a good guide on implementing Data Mesh in Databricks. There is plenty of content explaining what Data Mesh is, its principles, domains, data products, etc. But when you actually try to build it, the questions are much more practical:

  • How should you organize catalogs and schemas?
  • What should the central Data Team own?
  • What should domain teams be allowed to do?
  • How do you handle Unity Catalog, compute, CI/CD and monitoring without creating a mess?

Most of what I found was either very theoretical or covered only one small part of the implementation.

So I ended up putting together the guide I was originally looking for - Databricks Data Mesh Best Practices: A Practical Implementation Guide

u/Significant-Guest-14 — 10 days ago

If you have a Databricks MVP or Champion recognition, what has been the biggest benefit?

I've noticed that most discussions about Databricks recognitions focus on how to earn them. I'd like to talk about something else. What changed after you earned one?

In my case, becoming a Databricks MVP actually helped me earn the Solutions Architect Champion recognition. The community contributions recognized by the MVP program already satisfied several Champion requirements, making the journey much shorter than if I had started from scratch.

That made me wonder how others see the value of these recognitions. For those who have earned MVP, Champion, or both:

  • What has been the biggest benefit?
  • Has it helped your career or consulting work?
  • Has it increased trust with customers?
  • Did it open new opportunities?
  • Or was the biggest value simply becoming part of the community?

I'm collecting real experiences—both positive and negative—for an article comparing these recognitions from a practical perspective rather than just listing their requirements.

I'd really appreciate hearing your story.

P.S. If you'd prefer to continue the discussion on LinkedIn, or share a longer story there, feel free to join the conversation: [LinkedIn link]

u/Significant-Guest-14 — 16 days ago
▲ 42 r/databricks+1 crossposts

Keeping track of Databricks feature status (Preview → GA) is harder than it should be

One thing I've noticed is that I spend way too much time answering (or searching for answers to) questions like:

  • Is this feature GA yet?
  • Is it still in Public Preview?
  • When did it become GA?
  • What is the current name?

The information exists, but it's scattered across release notes, docs, blogs, and old posts.

A few days ago I shared an open-source side project that tracks Databricks feature renames. After reading the feedback here, I'm thinking that tracking feature lifecycle might actually be even more useful than tracking renames.

The idea would be something like this:

  • Public Preview → Beta → GA timeline
  • Rename history (if applicable)
  • Links to the official documentation
  • Dates when statuses changed
  • Eventually, the ability to follow a feature and get notified when something changes

Before spending time building it, I wanted to ask the community:

  • Would you actually use something like this?
  • What information about Databricks features do you find hardest to keep track of?
  • Are there other lifecycle events worth tracking besides Preview/GA?

If anyone is curious, the rename tracker that started this discussion is REbricked. I'm mostly interested in feedback on whether this direction solves a real problem.

u/Significant-Guest-14 — 28 days ago
▲ 97 r/databricks+1 crossposts

Lost track of Databricks product renames? Here's a community tracker

I recently contributed to REbricked, a community project that tracks Databricks product renames.

It includes:

  • Product rename history
  • Links to official documentation
  • A short quiz to test yourself

https://rebricked.org

We're still improving it, so feedback is very welcome. If we missed a rename or you have ideas, let us know.

u/Significant-Guest-14 — 1 month ago

Databricks Community Articles in Medium

Please follow https://medium.com/databrickscommunity for the best articles about Databricks.

We welcome contributions from anyone in the Databricks community. Submit your first article as a guest (follow us, and in the article menu, choose “submit to publication”). Accepted contributors may be added as regular writers.

u/Significant-Guest-14 — 1 month ago

Databricks AI Agent Genie Code is no longer Free. Now you have to pay as you go

On July 8, 2026, Databricks introduced one of the biggest changes to its AI experience.

Databricks AI Agent Genie products (Genie CodeGenie Spaces, and Genie One) are no longer free, as I wrote earlier. Instead, they now use a pay-as-you-go pricing model based on AI token consumption, with a free monthly allowance for each user before paid usage begins.

I wanted to understand how this works in practice. Not by reading the documentation. By using it.

>Also, be careful! Currently, only the Alert feature works, but the documentation states that there should also be a block after reaching the limit!
I think you need to notify users so they can use AI more carefully and wisely.

Full text - Medium

u/Significant-Guest-14 — 1 month ago

Merch from the Databricks Data and AI Summit 2026

While everyone is busy geeking out over the new features and major announcements from the Databricks Data + AI Summit, I decided to do a deep dive into what really matters: the vendor swag.

I did a quick analysis of the loot this year and realized a hilarious trend: vendors were giving away beanies, T-shirts, hoodies, socks, and even Crocs. You could literally build a complete wardrobe from scratch on the expo floor. The only thing missing? Pants, shorts, or underwear.

Jokes aside, the vendors went crazy this year. Beyond the usual sticker spam and keychains, the raffles were actually insane - LEGO sets, Nintendo Switch 2, drones, JBL speakers, and Apple gear were everywhere.

I decided to visualize this collection and put together two complete summer and winter looks.

I know I missed a ton of booths. I saw some people walking around with cowboy hats, bucket hats, bandanas, and a bunch of other random stuff.

What was the best/weirdest swag you managed to snag this year? Did any of you actually win the big raffles? Let me know what I missed!

u/Significant-Guest-14 — 2 months ago

What does the Databricks main office actually look like inside?

Got to visit the San Francisco HQ after the Data + AI Summit last week. Not open to the public — Databricks MVP status got me in. Figured some of you might be curious what it's actually like inside.

Quick rundown:

The office feels like the product. No unnecessary flash, everything has a reason.

Food ordering system — employees order individually through an app, each gets their own delivery. No cafeteria.

Dedicated bike room — not a hook on a wall, an actual room. SF culture is fully absorbed.

Free merch stand for visitors — great idea in theory. After DAIS, it was completely wiped out. Showed up too late.

Ice cream machine — didn't get to try it. Still thinking about it.

Tried to recreate the famous balcony photo. The view is legitimately incredible. Slightly terrifying height, though.

From the same balcony, you can see Meta, Salesforce, Anthropic, and LinkedIn offices across the street. Would love to visit those too — unfortunately, I don't know anyone there. Yet.

u/Significant-Guest-14 — 2 months ago

a well-known reporter at the DAIS

Does anyone know who this is? The only reporter at the DAIS with a personal bodyguard.

u/Significant-Guest-14 — 2 months ago

Databricks Apps nearly scale to zero (Cut Databricks Apps Costs by 76%: Automate Start/Stop)

Databricks Apps just... run. All the time. There's no scale-to-zero, so even if nobody opens your app for a week, you're still paying the full ~$350/mo (720 hrs). The thing is, most of our apps are internal dashboards and admin tools that people touch maybe a few hours a day. So we were basically paying full price to use ~15% of it.

What actually made me fix this: someone spun up a test app, forgot about it, and it sat there running for 69 days before we noticed on the bill. ~$800 for literally nothing.

The fix is kind of dumb, but it works great — two scheduled jobs that hit the Apps REST API:

  • A notebook that wraps the start/stop endpoints (takes app_name and app_command, plus an all option if you want to hit every app at once)
  • One job starts in the morning: 0 0 9 ? * MON-FRI *
  • One job stops it in the evening: 0 0 18 ? * MON-FRI *

That's 50 hrs/week instead of 168, so roughly 76% off. And honestly, nobody noticed — the app's just there during work hours.

Full text and example Notebook

u/Significant-Guest-14 — 2 months ago
▲ 6 r/DuckDB

I built a fully client-side quiz/testing app on DuckDB-WASM. The database lives in the browser, no backend at

I am glad to share another pet project. I kept running into study/quiz tools that force you to create an account and store everything on their servers. I wanted the opposite: a self-testing app where the data never leaves the browser. DuckDB-WASM turned out to be a perfect fit, so I built BOX-tests around it.

How DuckDB is used:

  • The entire app state: tests, questions, attempts, stats, lives in a single in-browser DuckDB instance. No server, no API, no accounts.
  • Persistence is just export/import of the .duckdb file (plus a JSON export for portability). "Saving" your data = downloading your database; "loading" = opening it back. Sharing a test = sending someone a file.
  • Analytics (progress, scores per group/difficulty) are plain SQL queries run locally — instant, no round-trips.

Things I ran into / curious about:

  • Cold-start bundle size of the WASM build and how aggressively people trim it.
  • Best practices for persisting the DB across sessions — right now I lean on file export + browser cache; curious whether folks here use OPFS / IndexedDB-backed persistence in production.
  • Whether anyone has patterns for schema migrations on a DB file the user holds (since I can't run migrations server-side).

Live (no signup, runs entirely in your browser): https://boxtests.com

Full disclosure: this is my own project, still v1. Sharing it here because the DuckDB-WASM angle is the actually-interesting part, and I'd love feedback from people who've pushed WASM persistence further than I have.

u/Significant-Guest-14 — 3 months ago

How do you monitor dozens of jobs in Databricks? I created free Databricks App

I run a fair number of scheduled jobs on Databricks, and monitoring them at scale always annoyed me. The two native options both fall short once you have a few dozen jobs:

  • Email/Teams alerts are easy to forget to set up on new jobs, so failures go silent.
  • The built-in Monitoring dashboard only shows the last 5 runs and spreads everything across many pages.

So I built a small Streamlit app that runs as a Databricks App and pulls everything through the Jobs API (via the SDK)

GitHub Repo

This is the first version. If you have any ideas on what to add, I'd be glad to hear them.
Here is the description of the application

u/Significant-Guest-14 — 3 months ago

How do you choose cluster and node types?

Databricks experts, I need your advice! How do you choose cluster and node types?

I've run numerous experiments on different cluster and node types, compared time and cost, and got some interesting results. I'll soon share my findings in a new article. This will help you save money or get better performance for the same price.

In Azure, Databricks offers over 300 node types. How do you choose from such a huge number of virtual machines? In practice, I have often seen that they choose among the first from the list:

  • General purpose (Delta cache accelerated) - 29
  • General purpose - 82
  • General Purpose (HDD) - 7
  • Memory optimized (Delta cache accelerated) - 41
  • Memory optimized - 79
  • Memory optimized (remote HDD) - 3
  • Storage optimized (Delta cache accelerated) - 17
  • Compute optimized - 18
  • GPU accelerated - 17
  • Confidential - General purpose - 4
  • Confidential - Memory optimized - 3
  • Confidential - Memory optimized (Delta cache accelerated) - 10
reddit.com
u/Significant-Guest-14 — 3 months ago

Modular structure for Databricks Apps (Streamlit)

Hey, I wanted to share something that's been bugging me for a while and get your take.

The official Databricks Streamlit tutorial puts everything into a single app.py. Fine for a demo. But the moment a real internal app grows past ~500–600 lines, it stops being fun:

  • Two people on the team touch the same file → merge conflicts every PR.
  • Hard to write unit tests when UI, data access, and business logic live in one module.
  • Git diffs become unreadable, and code review suffers.
  • When I point Cursor/Claude at the repo, it has to re-read the whole monolith on every prompt. Context window and cost both balloon.

So I refactored our internal template into something more boring and modular:

   app. py            # entry point only, routing
   pages/
   ├── home. py
   ├── analytics. py
   └── settings. py
   components/       # reusable UI bits
   services/         # SQL warehouse / UC / SDK calls
   assets/
   ├── styles.css
   └── logo.png
   tests/

This is my own repo, not a product. Sharing because the single-file pattern bit us hard, and I figured others might find it useful - https://github.com/protmaks/databricks_apps_streamlit_mod_template

u/Significant-Guest-14 — 3 months ago

I originally thought the Lovable + Databricks integration was kind of pointless.

Then I had a hackathon project where all the core work was already in Databricks: data processing, enrichment, and some ML. The only missing piece was a simple interface that non-technical users could actually open and use.

I tried Lovable mostly out of curiosity, and it worked better than I expected for an MVP.

A few practical things I learned:

  • The service principal needs access not just to the data, but also to the SQL warehouse / compute
  • I got it working on Databricks Free Edition
  • If every interaction hits Databricks directly, caching matters, and costs can add up pretty quickly

I still wouldn’t treat this as a production setup, but for a hackathon demo or fast idea validation, it felt surprisingly useful. I wrote more details in the article - https://medium.com/dbsql-sme-engineering/databricks-lovable-a-practical-case-study-and-what-it-costs-to-build-an-app-085f61b07126

u/Significant-Guest-14 — 4 months ago

I originally thought the Lovable/Databricks connector was kind of a gimmick.

Then I had a hackathon project where all the heavy lifting was in Databricks (data processing, enrichment, a bit of ML), but the result had to be shown as a simple app for non-technical users.

Tried Lovable mostly out of curiosity, and honestly, it worked better than I expected for an MVP.

A couple of practical notes in case anyone else tests it:

  • service principal needs access not just to the data, but also to the SQL warehouse / compute
  • I got it working fine on Databricks Free Edition
  • if you don’t cache responses, repeated queries can get expensive fast because you’re paying for warehouse runtime

I still wouldn’t treat this as my default production setup, but for demos / internal prototypes/idea validation, it was surprisingly useful.

I wrote a short article with examples - https://medium.com/@protmaks/databricks-lovable-a-practical-case-study-and-what-it-costs-to-build-an-app-085f61b07126

u/Significant-Guest-14 — 4 months ago