Stop Waiting for Spark: Enabling High Concurrency Mode in Fabric Notebooks
Last Updated on October 6, 2026 by Editorial Team
Author(s): Sandip Palit
Originally published on Towards AI.
As we scale our enterprise analytics, we constantly seek ways to optimize our data engineering workflows. We build elegant data pipelines, meticulously craft our transformation logic, and orchestrate complex workflows. Yet, despite our best efforts in code optimization, we frequently encounter a silent bottleneck that drains our time and our compute budget: cluster initialization latency. We are delighted to share a transformative approach to this challenge. In this detailed exploration, we will dive into High Concurrency mode in Microsoft Fabric Notebooks, a feature that fundamentally redefines how we manage Spark compute. By transitioning to this optimized execution model, we will slash pipeline latency from minutes to mere seconds while maximizing our Capacity Unit (CU) investments.

The Problem of Spark Cold Starts and Pipeline Latency
To appreciate the profound impact of High Concurrency mode, we must first examine the mechanical realities of standard Spark execution. When we orchestrate an enterprise-grade data platform, we typically break our logic into modular, domain-specific notebooks. We might have one notebook cleansing sales data, another aggregating marketing metrics, and a third calculating financial forecasts. While this modularity is excellent for code maintainability, it introduces severe friction when orchestrated in a traditional pipeline.
When we initiate a PySpark notebook for a data load or machine learning task, we are subjected to session initialization delays, frequently referred to as cold starts. There is an inherent latency in allocating containers, loading extensive libraries, and establishing the distributed environment. In a standard configuration, executing three separate notebooks in parallel forces the underlying engine to request, provision, and spin up three distinct Spark clusters.
This hardware allocation process takes time. It is not uncommon for a cold start to add two to three minutes of pure overhead to every single pipeline run. If we are running a micro-batch architecture that triggers every fifteen minutes, spending three minutes waiting for infrastructure allocation on every execution is mathematically unsustainable. Over the course of a day, we lose hours to idle waiting. Furthermore, we are not simply losing time; we are actively burning our provisioned compute budget just to turn the servers on. We clearly needed a paradigm shift to eliminate this redundant overhead, and High Concurrency mode provides exactly that solution.
How High Concurrency Works: Session Isolation vs. Cluster Sharing
To resolve the latency issue, we must change how our notebooks interact with the underlying compute pool. Historically, the relationship between a notebook and a Spark cluster was strictly one-to-one. High Concurrency mode breaks this rigid coupling, introducing a highly efficient one-to-many architecture.
When we enable High Concurrency, we instruct Microsoft Fabric to maintain a single, robust, and active Spark session. As our orchestration pipeline triggers subsequent notebooks, they no longer request new hardware. Instead, they seamlessly attach themselves to the already-running Spark session. Because the containers are already allocated, the libraries are already loaded, and the JVMs are already warm, the execution latency drops from several minutes to a fraction of a second.
The brilliance of this architecture lies in how it balances sharing with safety. We are sharing the heavy underlying compute infrastructure (the nodes, the memory pool, and the CPU cores), but we are isolating the execution contexts. The system intelligently multiplexes the incoming notebook commands, allowing multiple streams of code to process simultaneously on the same hardware without tripping over one another. We achieve the parallel processing capabilities of a massive, multi-cluster deployment while only paying for the footprint of a single, highly optimized cluster.
Enabling the Feature: Workspace Settings and Pipeline Activity Checkboxes
Implementing High Concurrency mode requires a deliberate, two-step configuration process. We must first establish the capability at the workspace level, and then we must explicitly invoke it within our orchestration pipelines.
First, we navigate to our Microsoft Fabric Workspace settings. Under the Data Engineering section, we locate our Spark Compute configurations. Here, we must define a custom Spark pool and explicitly enable High Concurrency capabilities for that pool. We take great care during this phase to appropriately size our nodes, ensuring the shared cluster has sufficient memory and core capacity to handle the aggregate workload of multiple notebooks running simultaneously. We also configure our timeout settings thoughtfully, allowing the shared cluster to gracefully spin down when all concurrent activities have successfully completed.
Once the workspace is prepared, we transition to our Data Factory orchestration pipelines. When we drag a Notebook Activity onto our pipeline canvas, we navigate to the settings pane for that specific activity. We will immediately notice a dedicated checkbox labeled “High Concurrency.” We must actively check this box for every notebook activity that we wish to route into the shared session pool. If we leave it unchecked, the pipeline will revert to its default behavior, spinning up a costly, isolated cluster for that specific task. By methodically enabling this across our pipeline architecture, we elegantly force our workloads to share the warmed compute resources.
Security & Variable Isolation: Preventing Data Leakage Between Shared Sessions
When we first introduce the concept of cluster sharing to our data engineering teams, a valid and critical question immediately arises: “If my Sales Transformation notebook and my Finance Transformation notebook are running on the exact same Spark cluster at the exact same time, will my variables clash? Will we experience cross-contamination of data?”
We can confidently assure our teams that Microsoft Fabric has engineered a robust defense against this scenario. The secret lies in strict REPL (Read-Evaluate-Print Loop) isolation. While the notebooks share the physical hardware and the overarching Spark application, each notebook is granted its own completely isolated Spark session and isolated namespace.
If we define a DataFrame named df_cleansed in our marketing notebook, and we simultaneously define a DataFrame named df_cleansed in our finance notebook, the underlying engine treats them as entirely distinct entities. There is zero risk of the finance data overwriting the marketing data in memory. Furthermore, temporary views created in one notebook session remain completely invisible to the concurrent sessions sharing the cluster. This airtight session isolation guarantees that we can safely execute highly sensitive, disparate workloads side-by-side without ever compromising our data integrity or enterprise security postures.
Cost ROI: Benchmarking CU Usage Before and After High Concurrency
Beyond the massive improvements in execution speed, the most compelling argument for adopting High Concurrency mode is the profound impact on our financial operations (FinOps). To fully grasp this return on investment, we must deeply analyze how Microsoft Fabric handles billing.
When we purchase a Microsoft Fabric capacity (for instance, an F64 SKU), we are provisioning a pool of 64 Compute Units. Every action we take within the platform, whether we are executing a complex query or running a PySpark model, consumes a fraction of these CUs. Fabric intelligently manages this through a system of bursting and smoothing. When we trigger a massive data engineering pipeline, Fabric automatically “bursts” to borrow compute power from the future to complete the task quickly. It then “smooths” this intensive usage by averaging the consumed CUs over specific timeframes, such as a rolling 24-hour window for background operations.
Before implementing High Concurrency, a pipeline triggering ten parallel notebooks would burst aggressively, requesting ten separate clusters. The sheer overhead of provisioning and tearing down these clusters consumed an enormous amount of our available Compute Units. If we ran this unoptimized architecture frequently, our smoothed average would rapidly breach our 64 CU limit, leading to severe platform throttling and degraded experiences for our business users.
By transitioning to High Concurrency, we radically flatten our consumption curve. We pay the initialization tax exactly once. As the subsequent nine notebooks attach to the active session, our CU consumption remains remarkably stable. When we benchmarked our daily pipeline executions, we observed a dramatic reduction in our overall smoothed CU consumption, often saving thousands of dollars in hidden compute overhead. This optimization allows us to run more jobs, process more data, and serve more users without ever needing to upgrade to a more expensive capacity tier.
Code-Based Demo: Optimizing Execution Under High Concurrency
To illustrate how we structure our internal execution logic, we rely on the mssparkutils library. This native utility is vital for passing parameters and triggering downstream processes. Below, we demonstrate how we define a pool of tasks that will benefit immensely from our shared compute strategy.
# Fabric Notebook logic optimizing execution under High Concurrency
# mssparkutils helps manage session isolation and concurrent notebook calls
from notebookutils import mssparkutils
# Execute multiple notebooks in parallel using the same active Spark session pool
# This avoids spinning up 3 separate clusters, saving massive CU costs
execution_pool = [
{"path": "/Notebooks/Transform_Sales", "params": {"Region": "NA"}},
{"path": "/Notebooks/Transform_Marketing", "params": {"Campaign": "Q3"}},
{"path": "/Notebooks/Transform_Finance", "params": {"Year": "2026"}}
]
# Note: In an orchestration pipeline, check "High Concurrency" in the Notebook Activity settings.
for job in execution_pool:
print(f"Triggering {job['path']} in shared session...")
# mssparkutils.notebook.run() executes sequentially;
# use Pipeline activities for true parallel shared execution.
mssparkutils.notebook.run(job["path"], 90, job["params"])
In this demonstration, we define our execution_pool dynamically, allowing us to pass specific, isolated parameters to each notebook. It is important to note our architectural methodology here: while mssparkutils.notebook.run() is incredibly powerful for chaining notebooks within code, it inherently processes sequentially. To achieve the true parallel execution described in this article, we take this exact modular logic and map it into Fabric Pipeline Notebook Activities, checking the High Concurrency box for each. This guarantees that all three domains, Sales, Marketing, and Finance, transform simultaneously on the single, warm cluster.
Conclusion
As we continue to mature our enterprise data platforms, our focus must shift from simply making things work to making things work efficiently. The days of accepting three-minute delays for every pipeline node are firmly behind us. By embracing High Concurrency mode in Microsoft Fabric Notebooks, we successfully harmonize velocity with fiscal responsibility.
We drastically reduce our end-to-end processing times, ensuring our business stakeholders receive their critical insights faster than ever before. Simultaneously, we protect our provisioned compute capacity, eliminating the wasteful overhead of redundant cluster initializations and safeguarding our environment against throttling. We highly encourage our fellow engineering teams to audit their existing orchestration pipelines, enable High Concurrency on their shared Spark pools, and experience the profound performance and cost benefits of a truly optimized data architecture. We will continue to explore and implement these brilliant infrastructure optimizations, ensuring our analytical foundations remain as resilient, secure, and lightning-fast as possible.
Hey, I am Sandip Palit, from Kolkata, India. I love to explore what’s new in the Data Science space and share it with the community. I am a Fabric Super User, and in this Microsoft Fabric Playlist, I will share my learnings on Microsoft Fabric and the tips and tricks of using it effectively..
Thank You for reading this article. Please feel free to share your thoughts in the comments section, and give this article a 🌟.
Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.
Published via Towards AI
Towards AI Academy
We Build Enterprise-Grade AI. We'll Teach You to Master It Too.
15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.
Start free — no commitment:
→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day
→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages
Our courses:
→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.
→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.
→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.
Note: Article content contains the views of the contributing authors and not Towards AI.