Attach files in Genie Code

Attach files in Genie Code

Genie Code is getting a new feature 🥳
Good news you can attach a document to Genie Code when you need it as temporary context for the current conversation.

What's a good use case?

🔴 Explain or summarize a document: “Summarize this PDF and list the key decisions.”

🔴 Extract information from PDFs, scans, or images.

🔴 Generate code based on a specification: attach a requirements document and ask Genie Code to create Python or SQL.

🔴 Analyze data files such as CSV, Excel, JSON, or JSONL.

🔴 Convert documentation into code, tests, or a notebook.

🔴 Interpret diagrams, charts, architecture sketches, or screenshots.

🔴 Migrate Tableau or Power BI files into AI/BI dashboards using /importBI.

What are the supported file types ?

.pdf, .csv, .xlsx, .xls, .json, .jsonl, and common image formats.

u/Youssef_Mrini — 23 hours ago

Getting ready for Genie Ontology

Are you getting ready for Genie Ontology?

You can leverage PAGES that sit in the Discover page and are organized by domain and subdomain.

Each domain and subdomain has its own set of Pages and users with access to a domain can create and govern them.

🛑 But first, what do you mean by Pages?

Pages are part of UC semantics; it's the business context that you define and govern explicitly, forming the human-modeled layer of the Genie Ontology.

🛑 Why is it useful?

When Genie One answers a question about a concept you've defined in a Page, it prioritizes the Page's definition over context it infers automatically, and cites the Page so users can confirm the source.

🛑 Any tips to build pages?

You can create Pages from those documents instead of writing each one by hand. Genie Code reads the documents you attach, extracts the terms it finds, and returns a set of proposed Pages. You review and edit the proposed Pages before any of them are created.

🛑 Is it a collaborative environment?

You can Comment: Ask a follow-up question or flag context for the owner.

You can Suggest edits: Suggest changes to a published Page's body.

Each time you click Suggestion, edit the body, and click Save, your edits are grouped into a single batch.

The owner or curator accepts or rejects the entire batch at once. Accepting a batch clears all other pending batches on the Page, including those from other users, and this can't be undone.

You can React: Upvote or downvote a Page to signal whether it answered your question.

The owner or curator can also edit a published Page's content directly, bypassing the suggestion workflow.

🛑What's next?

Create domains, Subdomains, leverage UC metric views, and connect your external tools to Databricks

u/Youssef_Mrini — 7 days ago

🔴 Unity AI Gateway is Generally Available. 🔴

Unity AI Gateway is the Databricks governance solution for AI and is part of Unity Catalog.

You can:

⚡ Control which AI services teams can use.

⚡ Route and manage AI traffic across providers.

⚡ Govern MCP servers to control access and costs.

⚡ Monitor usage, cost, access, and lineage from one place.

FYI: Some capabilities, including service policies and agent services, remain in Beta.

Unity AI Gateway Documentation: https://docs.databricks.com/aws/en/ai-gateway

Blog post : https://www.databricks.com/blog/unity-ai-gateway-generally-available

u/Youssef_Mrini — 16 days ago

What's new on Genie One July 2026 ?

AI Capabilities & Memory

  • Genie Code Mentions: Type u/Genie Code in a Notebook comment to get help directly within the comment thread.
  • Copilot Integration: Connect Genie directly to Microsoft Copilot Cowork. 📖Documentation
  • Memory (Beta): Genie One can save specified details (like preferences and workflows) to apply in future conversations. 📖Documentation
  • Recall & Reuse Context: Search past conversations and bring that context into your current chat. 📖Documentation

Document Features

  • Document Links: Documents supports embedded hyperlinks.📖Documentation
  • Edit Timestamps: Displays the exact date and time documents were last edited. 📖Documentation
  • Document Citations: Documents display citations linking content directly back to source data.
  • Version History: Review earlier versions of documents through built-in history tracking.
  • Soft Tabs: Open document links in soft tabs without navigating away from the current document.

Integrations & MCP Apps

  • Genie One MCP App (Beta): Adds UI interactions from third-party agents, including interactive visualizations and Genie Ontology citations. 📖Documentation
  • GitHub MCP (Beta): GitHub MCP connection for Genie One is available📖 Documentation.

Scheduled Tasks

  • Shared Task Overview: Access a dedicated modal and listing page to view scheduled tasks shared with you by others.
  • Task Mentions: u/mention scheduled tasks directly inside Genie One.
  • On-Demand Execution: Select Run now while authoring a scheduled task to run it immediately.

Platform, Data & UI Enhancements

  • Ontology Snippets (Public Preview): Support added for ontology snippets (contact your Databricks account team to enroll). 📖Documentation
  • Partially Supported Datasets: Ask Genie now queries valid datasets rather than failing completely when a dashboard contains unsupported SQL expressions.
  • Front-end Private Link: Configure account-level front-end Private Links by turning on the Custom URLs and Account preview.
  • Certified Visualizations: SQL charts and visualizations inserted by Genie One are now read-only certified visuals.
  • Inline Conversation Renaming: Edit conversation titles inline by clicking the title in the chat header.
  • Mobile Workspace Sign-in: Enter your workspace URL manually in the Genie mobile app if automatic discovery fails. 📖Documentation
  • Expanded PDF Limit: The maximum character limit per uploaded PDF increased from 4,000 to 15,000 characters. 📖Documentation
u/Youssef_Mrini — 16 days ago

What's new in Genie Agent July 2026

  • Prompts and responses initiated by Genie Chat are visible in Genie Agents monitoring when the sharing is enabled.
  • You can upload local Excel and CSV files to an Agent Mode chat for ad-hoc analysis. 📖 Documentation
  • You can retrieve visualization results from Genie Agents using the API. 📖Documentation
  • You can save a Genie One thread as a Genie Agent so you can open it in your workspace to fine-tune its context, instructions and data assets.

  • You can share individual Genie conversations to collaborate with teammates or allow agent managers to review your results.
  • Genie Agents can analyze files stored in Unity Catalog volumes such as PDFs, slide decks, and images. 📖Documentation
  • Agent mode APIs for Genie Agents are available. 📖Documentation
  • You can download visualization results through the Genie Conversation API on Private Link workspaces in supported regions.
  • Databricks Genie App for Teams: 📖Documentation
  • Databricks Genie App for Slack: 📖Documentation
u/Youssef_Mrini — 17 days ago
▲ 0 r/cursor

Databricks has officially launched its verified plugin on the Cursor Marketplace

If you're building with Databricks and coding in cursor, this integration brings a powerful suite of Databricks capabilities directly into your AI-assisted workflow.

Here is a breakdown of what the plugin includes:

🔹 10 Specialized Skills

🔹 Smart Routing Rules

🔹 Built-in Commands

This plugin bridges the gap between your local IDE and your data workflows.

Check it out on the Cursor Marketplace today! 👉 https://cursor.com/marketplace/databricks

reddit.com
u/Youssef_Mrini — 23 days ago

Databricks has officially launched its verified plugin on the Cursor Marketplace

If you're building with Databricks and coding in cursor, this integration brings a powerful suite of Databricks capabilities directly into your AI-assisted workflow.

Here is a breakdown of what the plugin includes:

🔹 10 Specialized Skills

🔹 Smart Routing Rules

🔹 Built-in Commands

This plugin bridges the gap between your local IDE and your data workflows.

Check it out on the Cursor Marketplace today! 👉 https://cursor.com/marketplace/databricks

reddit.com
u/Youssef_Mrini — 23 days ago

Unit testing for pipelines

https://preview.redd.it/pq5958bepxfh1.png?width=2446&format=png&auto=webp&s=d5ef8532fa430ffb470d9c4d81f5339f297a0ba4

Lakeflow Spark Pipelines Editor now supports native Python unit testing, letting you validate your Python and SQL transformation logic right where you build it.

By using mock data, you can safely pressure test edge cases, iterate on table-identifier operations, and verify proprietary pipeline APIs including Auto CDC, streaming tables, expectations, and append flows.

reddit.com
u/Youssef_Mrini — 23 days ago

RBAC in Databricks

https://preview.redd.it/fcv2rg4wwxeh1.png?width=1015&format=png&auto=webp&s=99b7598bb5ce999ff5171b3ae29dd626b8fb82d2

Role-based access control (RBAC) lets users assume a role in Databricks, using only that role's permissions for the duration of the session.

RBAC enables role-based access control: users must assume a role to access sensitive data, preventing them from accessing it when acting as their user identity and from mixing data across use cases.

To learn more about RBAC: https://docs.databricks.com/aws/en/security/auth/rbac/

reddit.com
u/Youssef_Mrini — 28 days ago

Genie cost tracking

🔴 Genie Cost tracking update🔴

You can now track GENIE_FREE_USAGE SKU (only starts appearing on July 20, 2026)

FYI: Free usage consumed before this date is not visible in the system tables.

All free Genie usage appears under sku_name = 'GENIE_FREE_USAGE' it does include usage under the free allowance( 150 DBUs) which resets on the first of each month.

Until 🔴 July 31, 2026🔴 , all Genie One and Genie Agents usage is free, and captured under the GENIE_FREE_USAGE SKU.

To distinguish between products within the free usage SKU you can filter on usage_metadata.genie.surface:

GENIE_CODE: Genie Code free usage

GENIE_ONE: Genie One

GENIE_AGENTS: Genie Agents

This free usage SKU tracks consumption but deliberately has no list price entry in the system tables. Because it is completely free, joining the usage and price tables will naturally return no match for this item.

u/Youssef_Mrini — 29 days ago

Google Gemini flash 3.6 is already available

Google’s Gemini 3.6 Flash is already available on Databricks. Yes that's true🎉

Always remember the 4C:

🙏Control: Governance over the multi modal multi agent estate. The anchor is Unity AI Gateway and Omnigent for agent orchestration.

🙏Choice: You get to choose between the different models available on the platform.

🙏Context: The anchor is Genie Ontology the self improving knowledge graph that automatically extracts business meaning from the assets.

🙏Cost: The anchor is Unity AI Gateway which adds hard spend caps to stop runaway agent costs. Omnigent's cost control plays the same role at the coding agent level.

u/Youssef_Mrini — 30 days ago

Genie Ontology

https://preview.redd.it/cnkpxe9f6keh1.jpg?width=1260&format=pjpg&auto=webp&s=4ce68cad93912763ffba1956eb690f1f6940552a

Genie Ontology is coming soon. You can already get ready by preparing the following components:
🛑UC Metric Views: Build measures, dimensions, and relationships that map to your business KPIs with AI-assisted authoring
🛑Glossary: Define authoritative concepts and taxonomies for agents and humans to reason over data
🛑Domains and subdomains: Organize assets into business-aligned groups for scoped, relevant and certified context 
🛑Genie Knowledge Store: Automatically build a permission-aware context graph that captures knowledge at scale

reddit.com
u/Youssef_Mrini — 1 month ago

What’s new in Databricks Data+AI Summit Edition Part 1

The scale of what was accomplished over those few days was historic breaking records across the board:

🛑 100K Total Attendees: A powerhouse community of data minds, including 32,000 innovators packing the venue in person, with the rest of the global community tuning in remotely from every time zone.

🛑174+ Countries Represented: A truly global phenomenon, uniting diverse perspectives and top talent from every corner of the world.

🛑240+ Exhibitors: A large ecosystem showcase featuring the industry-leading partners and technologies shaping tomorrow’s data landscape.

The summit may be over but the momentum is just getting started. The connections made, skills learned and breakthroughs shared have officially set the trajectory for the next era of data and AI!

Read the full newsletter : Part 1

u/Youssef_Mrini — 1 month ago

Sponsored posts everywhere

Big thanks to LinkedIn for reducing post reach by 90% just to push sponsored content.

To make it worse, I keep getting random sponsored posts that aren't even related to my preferences

reddit.com
u/Youssef_Mrini — 1 month ago

What's new in Apache Spark 4.2 ?

Apache Spark 4.2 is already available on DBR 19 and it reflects the strength of the Apache Spark community with more than 1,900 commits from over 260 contributors

https://preview.redd.it/259r9wl9oldh1.jpg?width=941&format=pjpg&auto=webp&s=67aa094b8e2e7b6bbf044101b4fa876dd906ae0f

What's new in Apache Spark 4.2 ?

🛑 Metric views: a native semantic layer that lets you define governed business metrics once and reuse them consistently across SQL, BI, apps and AI, preserving correct aggregation semantics. ( Open source is on our DNA)

🛑Spark Connect & PySpark: better Spark Classic compatibility (RDD APIs, YARN cluster mode, error propagation) + easier remote invocation of Spark as a service.

🛑Arrow-first Python: Arrow-optimized Python UDFs on by default, Pandas 3 support, Arrow UDFs and zero-copy interop with Polars/DuckDB via the Arrow Data Interface

🛑AI-native SQL: vector similarity/distance functions and NEAREST BY top-K retrieval for search, recommendations and RAG-style workloads.

🛑Native geospatial: built-in GEOMETRY/GEOGRAPHY types and ST_* functions, with Parquet, WKT/WKB and SRID support.

🛑More SQL primitives: SYSTEM.BUILTIN/SYSTEM.SESSION qualification, SET PATH search paths, SQL cursors, QUALIFY, time_bucket, tuple sketches, top-K max_by/min_by and IGNORE/RESPECT NULLS.

🛑Auto CDC in Declarative Pipeline: first-class SCD Type 1 change-data processing via a Python API, replacing hand-written merge logic.

🛑Real-Time Mode for PySpark( Like the flash) millisecond-latency streaming now extended to stateless PySpark queries.

🛑Data Source V2 advances: first-class CDC via the CHANGES clause, MERGE INTO performance gains, schema evolution for INSERT INTO, and transaction API foundations.

🛑Platform improvements: modernized Web UI (Bootstrap 5, dark mode), better Kubernetes support and JDK 25.

reddit.com
u/Youssef_Mrini — 1 month ago

Table update triggers on OpenSharing and system tables

able update triggers just got a lot more interesting. 🚀

Until now Lakeflow table update triggers only fired on local tables in your own workspace. As of the new Beta they also work on:

→ OpenSharing tables (data shared into your workspace via Delta Sharing)

→ System tables (system.billing.usage, audit logs, lineage…)

Why is that a big deal technically ?

You can now run a job the moment shared data updates across workspace, regional & cloud boundaries. The trigger watches the shared table and fires when new data actually arrives.

You're turning Delta Sharing from a passive data-access layer into an active, event-driven integration layer.

A Lot of benefits

1-Less orchestration

2-lower latency

3-no idle scheduled runs.

Want to play with it? Docs + setup requirements here

https://docs.databricks.com/aws/en/jobs/trigger-table-update#trigger-on-ds-and-sys

u/Youssef_Mrini — 1 month ago

Govern AI Functions with Unity Catalog

You can now use Unity Catalog permissions to restrict which task-specific AI Functions your organization can access, independent of foundation model access.

In the example below, I'm giving access to all the users the permission to use the AI Functions ( AI_Classify, AI_Extract, AI_parse_document)

For more information: https://docs.databricks.com/aws/en/large-language-models/ai-functions-uc-permissions

Reach out to your Databricks account team to enable this feature.

Document Intelligence on Databricks: https://www.youtube.com/watch?v=sdG73gI143c

Get started with AI Functions: https://www.youtube.com/watch?v=Dt7K8As4Qh8

https://preview.redd.it/rjisx7iab0dh1.jpg?width=1238&format=pjpg&auto=webp&s=806a132b0077c7a7f88982cbef780fe6fab1c6dd

reddit.com
u/Youssef_Mrini — 1 month ago

Inside Lakehouse//RT and the Reyden Engine With Himanshu Raja Sr Director PM at Databricks

Lakehouse//RT was built for the specific demands of real-time serving at scale,

Millisecond latency, at any scale: Reyden’s fully asynchronous execution model delivers response times as low as 10 milliseconds on smaller datasets and 100 milliseconds on larger ones, without latency degrading as throughput climbs into the tens of thousands. And unlike engines optimized only for simple lookups, Lakehouse//RT applies state-of-the-art performance techniques to the full range of analytical complexity.

youtu.be
u/Youssef_Mrini — 1 month ago
▲ 29 r/databricks+1 crossposts

Databricks Medallion Architecture: Catalogs by domain or by Bronze/Silver/Gold?

I got into a debate at work today about how Unity Catalog should be organized, and I'm curious what others are doing in production.

Which approach do you prefer?

Option 1: Catalogs by medallion layer

bronze

silver

gold

Then have separate catalogs for curated data products or analytics built from the Gold layer.

Option 2: Catalogs by business domain

finance

operations

hr

Then each catalog contains schemas like bronze, silver, gold

What are the strongest reasons to choose one approach over the other?

reddit.com
u/TheManOfBromium — 1 month ago