Skip to content

writing / the stack report /

Towards Async Transactions in Django

async with transaction.atomic() is coming to Django, hopefully for 6.2. How we got there at Django on the Med 🏖️, and what the rules will be.

Concurrency — where multiple things can happen, in any order — doesn't combine well with database transactions, which are necessarily well-ordered.

Imagine two tasks sharing a single connection. The first opens a slow transaction, which eventually fails and is rolled back. Meanwhile the second makes a quick single write, that is lost when that transaction rolls back:

import asyncio

from django.db import transaction

from myapp.models import Log


async def slow_transaction(transaction_opened, quick_write_done):
    async with transaction.atomic():
        await Log.objects.acreate(msg="Example write within the transaction")

        # Mark slow transaction as begun, and
        # wait for quick write to finish.
        transaction_opened.set()
        await quick_write_done.wait()

        # An error here will roll back the transaction.
        raise RuntimeError("slow_transaction fails")


async def quick_write(transaction_opened, quick_write_done):
    # Wait until the slow transaction has begun.
    await transaction_opened.wait()

    await Log.objects.acreate(msg="A quick single write")

    # And mark that we're done.
    quick_write_done.set()


async def main():
    """
    Interleave slow_transaction and quick_write showing why transaction
    separation is required.
    """

    # Use asyncio.Events for sequencing.
    transaction_opened, quick_write_done = asyncio.Event(), asyncio.Event()

    try:
        async with asyncio.TaskGroup() as tg:
            tg.create_task(slow_transaction(transaction_opened, quick_write_done))
            tg.create_task(quick_write(transaction_opened, quick_write_done))
    except* RuntimeError:
        pass

    # The error in slow_transaction meant that quick_write's create was rolled
    # back.
    assert not await Log.objects.filter(msg="A quick single write").aexists()

The async with transaction.atomic() doesn't work yet — that's today's topic — but imagining it did, if concurrent (and not otherwise serialised) async tasks share a database connection, then transaction boundaries are broken.

This isn't a Django-specific problem. Fundamentally, database connections are serial pipes of requests to the database and responses from it. Queries must be well-ordered. A transaction is a series of queries starting with a BEGIN and ending (on the happy path) with a COMMIT. Every query in-between must belong to the transaction. Any query that does not belong to the transaction must wait until the transaction is complete.

Until now, Django has required using the ORM's sync API for transactions. You'd wrap your ORM code in sync_to_async, where it would be executed on a worker thread, synchronously, and no concurrency issues arise. (As soon as you start thinking about async transaction scenarios, ensuring that database queries are correctly serialised, and that transactions are sequenced, is the key feature of the existing approach.)

Other ORMs offering an async interface have to handle this same problem in their own way. They either force you to manage the transaction ownership explicitly, e.g. SQLAlchemy, or tie the transaction to the current task, e.g. Peewee. We're all shuffling the same set of constraints into more or less similar piles, as best meshes with the context we're coming from.

Nonetheless, adding the async with transaction.atomic() API to Django remains desirable. For standard async views, normally running as a single asyncio task, being able to group related ORM operations into a transaction, without needing to move away from async methods, would essentially complete the ORM's async API.

So, we sat down at Django on the Med 🏖️ to see if we could make progress. The good news is, we certainly did. I'm hopeful that we'll be able to get this in for Django 6.2. 🤞

§

Now don't panic. It's not a bait and switch. Having got you hooked, I'm not going to turn the rest of the piece into an advert for Django on the Med 🏖️.

To some extent, it can't help but be that, an advert. As you'll see, it's the sprints process working exactly as intended, as hoped. I've made no secret of how successful I think the sprint events are, and how powerful they can be in pushing Django forward. But, if you want to hear (or watch) more on that, I'll point you to the recent Django Chat episode on Django on the Med 🏖️, where Will and I discuss the recent event in Pescara, as well as the other news from over the summer.

§

Database connection management in Django is per-thread. When you access connections["my_db_alias"], each thread gets its own connection for the alias, and since a thread only does one thing at a time, its queries are naturally well-ordered.

Under asyncio, every task in an async application runs on the same thread, on the event loop, interleaving at each await. Multiple tasks using the same connection is exactly the problem we started with. For this reason (and because sync calls block the event loop) Django enforces that you can't use the ORM's sync API directly in async code, raising a SynchronousOnlyOperation exception if you try. The ORM's async API — aget(), acreate(), &co. — hands the actual work off with sync_to_async(…, thread_sensitive=True). Queries are run synchronously on a separate worker thread, with its own connection, whilst async tasks await the results.

The thread-sensitive argument to sync_to_async means that all ORM work will go to the same worker thread, serialised onto one database connection. Generally that's fine for a single HTTP request. It's what we're used to from sync Django — where a single worker, in a single thread, handles a single request at a time, with a single database connection.

With concurrent HTTP requests in play on the same event loop, though, we need to be able to provide multiple worker threads, for multiple connections. Otherwise every request would need to queue for the single database connection.

For this, asgiref.sync provides what's called a ThreadSensitiveContext. This allows sync_to_async to use a different worker thread (and, again, so a different connection) for each context. Django's ASGIHandler wraps each separate request in its own ThreadSensitiveContext:

class ASGIHandler(...):
    ...

    async def __call__(...):
        ...

        async with ThreadSensitiveContext():
            await self.handle(...)

This means that, in practice, it's one worker thread, and one connection, per request.

That's how things have worked for the last few years, and going into Django on the Med 🏖️, I had:

  • A sketch of adding async with transaction.atomic() which worked for the single-task case.
  • The foot-gun example from the beginning in hand.
  • And, an idea for nesting ThreadSensitiveContexts — which currently is a no-op — to allow new worker threads, so new connections.

What I didn't have was a clear idea of how to put these together.

§

On Day 1 of the sprints, after the intro session, I was able to gather an expert panel to help me sort through the difficulties. Lily Acorn, who's on the Steering Council, and is an experienced ORM contributor. Simon Charette, who's the leading expert on the ORM. Jacob Walls, who's a current Django Fellow. And Charlie Denton and Sam Searles-Bryant, who're from Kraken Tech, one of the largest users of the ORM at scale. Charlie and Sam are also the authors of the django-subatomic package, which is working its way into Django. They're experts on nested atomics, and transaction discipline. I really couldn't have had a better group to work with.

In an hour, we were able to work through the issues, see the difficult cases, back and forth over them, and come up with a clear picture of the constraints that async with transaction.atomic() would have to satisfy, if we're going to be able to add it to Django.

I cannot stress enough that this process was absolutely invaluable.

From that discussion, I was able to write a comment on the issue explaining those constraints: an async transaction would need its own connection — not that of the existing shared worker thread; hence, we would need to allow nesting ThreadSensitiveContexts to provide that connection, and, finally, we'd ensure that only the async task opening the transaction could drive transaction state and savepoints.

I finished the first morning fired up that it was going to be feasible, and we had an afternoon bike ride by the coast to unwind.

Some of the Django on the Med 🏖 riding bikes along the coast in Abruzzo

Credits to Paolo Melchiorre for the action shot of the ride 🚴

On Day 2 my plan was to knock off a review of a Django ticket about using aclosing() with async generators. It had been on my backlog for six months, and had a monster of a report, so I needed a good window. I was also working with Mark Smith on django-bgt which was taking shape nicely.

About mid-morning, Caleb Hattingh called me over. He'd been looking at the comment I'd left on the async with transaction.atomic() ticket the day before, and started to investigate.

Unbeknownst to me, Caleb is one of those rare folks who actually knows about async — it turns out he even wrote a book on it — so, not only was he interested, he got the proposal and saw where it was headed. This was an unexpected find. That afternoon and evening Caleb and I discussed the issues at length, including diversions into async and threading in general, in Rust, and Python, and beyond.

On Day 3, Caleb and I were able to discuss the asgiref changes to ThreadSensitiveContext. Caleb identified the need to handle complex cases where there's nesting of sync and async contexts (which is always a tripping point in this stuff). We got as far as being confident that the proposal, at least on the asgiref side, would work, but still needing some definitive test cases to show that with certainty.

This was the it's going to work moment. I left Pescara pretty sure that we'd be able to get async with transaction.atomic() into Django.

Caleb was obviously fired up too. The next weekend Caleb opened dual PRs on asgiref and on Django to implement this. The asgiref PR is essentially as we had it in Pescara, and is correct, and we're working the Django PR into shape.

From the user perspective it'll work how you expect. Entering an async with transaction.atomic() will create a nested ThreadSensitiveContext which gets a fresh worker thread, and so its own connection, for the transaction it manages. When the block exits, that connection is cleaned up, and we're back to the request's context, thread and connection:

async def my_view(request):
    await Log.objects.acreate(msg="before")

    # The atomic block creates a nested context, with a separate worker thread
    # and connection.
    async with transaction.atomic():
        await Log.objects.acreate(msg="inside")

    await Log.objects.acreate(msg="after")

If you remember the example from the beginning, we can run exactly the same code again, just flipping the final assertion:

# quick_write's create survived slow_transaction's rollback.
assert await Log.objects.filter(msg="A quick single write").aexists()

When slow_transaction enters atomic(), the nested ThreadSensitiveContext is set in its copy of the context. quick_write is a sibling task, created with its own context, so it never sees slow_transaction's connection. It carries on using the request's worker thread and connection, in autocommit, and its write is committed independently. The transaction boundary is maintained.

That's what's needed for the usual cases. Folks will try and do funny things, though, so it'll come with some additional rules:

First. Only the task that opened the atomic() block can nest further atomic() blocks inside it. Transactions must be well ordered. Asynchronous tasks are, in general, not. What's more, savepoints and nesting are complex, even in straight-line code. Once you enter an async atomic block, only that same task may make further use of atomic(). Try from any other task and it will raise a TransactionManagementError. This enforces that nested atomic() calls are well ordered.

Second. Avoid creating subtasks inside an atomic() block. A transaction is one sequential unit of work, so keep it in one task. The database connection is a serial pipe in any case. You gain little, bar complexity, by trying to make queries concurrent.

Third. If you must use subtasks, use structured concurrency. Use a TaskGroup (or an awaited gather()), so that every subtask has finished before the atomic() block exits. Their queries are serialised onto the transaction's worker thread, and so are part of the transaction. Anything still running after the block exits is not.

async with transaction.atomic():
    # OK (rule 3): both tasks complete, inside the transaction, before the block exits.
    async with asyncio.TaskGroup() as tg:
        tg.create_task(Log.objects.acreate(msg="one"))
        tg.create_task(Log.objects.acreate(msg="two"))

    async def nested():
        async with transaction.atomic():
            ...

    # Raises (rule 1): not the task that opened the transaction.
    await asyncio.create_task(nested())

Rule 1 is enforced by the TransactionManagementError exception.

Rule 2 is a guideline, to save yourself.

Rule 3 isn't enforced, yet. You have to ignore the advice of Rule 2 to get there, but it's a sharp edge, and I'd like Django to catch it before 6.2, if we can.

There are a couple of options here. On exit, the atomic() block could mark its context as closed, so any late query raises, rather than running outside the transaction. Or it could maybe keep track of the other tasks that have used the transaction, and complain if any are still pending when the block exits. (Maybe both: the first as the safety net, the second for a better error message.) But these are niceties. The reality is that most async Django views are not going to be making parallel queries, on multiple connections, with concurrent tasks. The core functionality is what matters, and looks like it will be available.

Using async with transaction.atomic() does entail an additional connection for the duration of the transaction. Connection pools will need to be tweaked appropriately. As per the existing advice, persistent connections using CONN_MAX_AGE should be disabled.

§

If you'd like to give this a play, try the Django PR against your own async code. Let us know how you go!

Huge thanks to Caleb, for picking this up and running with it, and to everyone who's helped along the way, and particularly Lily, Simon, Jacob, Charlie, and Sam for the discussion on Day 1 at Django on the Med 🏖️.

And — I guess it is an advert 🥳 — thanks to Paolo, for hosting the event, making the space where this kind of thing happens. 🎁

See you all next year!

The Stack Report goes out by email, approximately monthly. Subscribe