October 1, 2026
|

By Tal Zussman

TrieDB (also known as MonadDb), the database behind Monad, runs on a dedicated NVMe device rather than on a file system. Like most databases running directly on a block device, it has to decide how much of its I/O should go through the operating system kernel’s page cache, which stores recently accessed disk data in memory. In March 2025, Linux 6.14 added a per-I/O flag, RWF_DONTCACHE, that lets an application decide how data should be cached per I/O request rather than per file, giving applications much more fine-grained control over their data. We wanted to take advantage of this functionality for TrieDB, but unfortunately, RWF_DONTCACHE was not available for block devices.

We decided to implement support for RWF_DONTCACHE on block devices ourselves, which required modifying the kernel’s I/O completion path, and decided to upstream our changes. What we thought would be a small, straightforward change evolved into a more generic interface which has been merged into Linux 7.3 and is now being used by other parts of the kernel as well. At Category Labs, we’re dedicated to using, developing, and contributing back to open source software, across our entire stack.


TrieDB and the page cache

TrieDB manages its NVMe disk directly rather than through a file system, avoiding file system allocation, journaling, metadata, and locking, none of which the database needs.

However, it does not skip the page cache. The page cache stores recently read and written disk data in memory. Writes are buffered in the page cache and written to the device later through writeback. The cost is that the kernel decides what data is cached, using an approximate LRU (least recently used) policy. Any cold data that is read in and only used briefly can end up evicting the working set to make room for it, while a kernel thread (kswapd) spends CPU choosing what data to evict.

A common alternative is direct I/O. O_DIRECT bypasses the page cache, so nothing is evicted, but it applies to all I/O on the file descriptor. I/O requests must be aligned to the device block size. Each write goes to the disk separately, so small writes are not merged, and there is no readahead (the kernel mechanism through which file data is speculatively prefetched into the page cache). Mixing direct and buffered access to the same data is generally not recommended. And since the kernel no longer caches anything, the database typically has to implement its own cache in order to maintain performance. A block device opened without O_DIRECT uses the page cache exactly like a file does.

While TrieDB previously used exclusively direct I/O, it now opens its disk without O_DIRECT, so data is read and written through memory when possible. This has improved performance by allowing commits to finish faster before data is fully flushed to disk, and has also reduced disk accesses when handling RPC queries, as they typically query recent changes which are still present in the page cache.


What RWF_DONTCACHE does

By using the page cache, TrieDB inherits the kernel’s existing LRU eviction policy and is subject to its decisions. However, TrieDB may be able to better understand which data is more likely to be reused, resulting in a future cache hit, better performance, and less eviction overhead. A control mechanism, like RWF_DONTCACHE, would allow it to programmatically take advantage of that insight.

RWF_DONTCACHE is a flag for preadv2(), pwritev2(), and io_uring read and write requests. Jens Axboe proposed it in 2019 as RWF_UNCACHED, and it was merged in Linux 6.14 (LWN). It is a per-I/O flag rather than a per-descriptor one, so the same file descriptor can mix cached and uncached requests.

For a read, the kernel goes through the page cache as usual, copies the data to the user buffer, and then removes the pages from the page cache, unless they were already present: cached data stays cached. Only pages that were brought in for this read operation are dropped.

For a write, the kernel copies the data into page cache pages, marks them with a “dropbehind” flag, and starts writeback immediately rather than waiting. When the device reports the write complete, the completion path sees the dropbehind flag and removes the pages from the cache. Writes still merge in the page cache before submission, and the application does not wait for the device to complete the write.

With this approach, we can get the best of both worlds. Hot reads can use the cache normally, while cold reads and one-time writes use DONTCACHE. Because a DONTCACHE read leaves cached pages alone, it cannot evict the working set. TrieDB has both kinds of traffic. For example, execution repeatedly reads recent state (high cache hit rate), while some RPC queries read cold, historical data just once. Existing alternatives, like posix_fadvise(POSIX_FADV_DONTNEED), do not drop dirty pages, so for writes the application has to force writeback first with sync_file_range() or fsync() and call it again afterwards.

Each file system has to implement support for RWF_DONTCACHE individually. XFS did so in Linux 6.15, ext4 in 6.17, and the NFS client in 6.18. Before 7.3, block devices rejected the flag with EOPNOTSUPP.


Why block devices were different

At first, this seemed like a simple enough change. Block device operations go through the same generic page cache code as the file systems that support DONTCACHE, so accepting the flag is a one-line change to the block device file operations and a few more lines to check for the flag in the submission path, which is what our first patch did. While this works for reads, writes are a problem.

The write path: dropbehind folios are written into the page cache and packed into a bio for the NVMe device.
The write path: dropbehind folios are written into the page cache and packed into a bio for the NVMe device.

When the device finishes a write, it raises an interrupt, and the kernel runs a completion handler in interrupt context. Under this mode, the CPU executes on behalf of the device rather than any thread, and the handler cannot block or sleep. However, removing a page from the page cache requires taking the page’s lock and possibly unmapping it from processes that have it mapped, both of which are operations that can sleep. As such, the dropbehind removal might need to sleep while the write-completion handler runs in a context where sleeping cannot happen, so removal can only run in preemptible task context.

File systems that support DONTCACHE already deal with this. XFS, for example, passes off write completions to a per-inode workqueue because it may need to update metadata at completion (a blocking operation), and it routes DONTCACHE completions to that workqueue as well. Currently, block devices and some older file systems complete writes directly from the interrupt. The generic DONTCACHE code handles that case by checking whether the completion is in task context. If not, it simply skips the removal. With our initial simple change, a block device would have accepted the flag and left every written page in the cache.

Dropbehind folios reach interrupt-context completion, can’t be dropped there, and stay in the page cache.
Dropbehind folios reach interrupt-context completion, can’t be dropped there, and stay in the page cache.


Deferring completion in the block layer

Based on feedback from some Linux maintainers, our next versions deferred the page removal itself: when the completion ran in interrupt context, it appended the page to a page cache-specific per-CPU list and scheduled a kernel worker to remove it later. However, Matthew Wilcox, the page cache maintainer, pointed out that this functionality was being added for one completion use case when several other completion handlers have the same need for different reasons. For example, checksum verification, decompression, and error reporting all need task context at completion time, and each subsystem that does them has been maintaining its own workqueue to get it. He suggested a more general interface to unify these cases instead.

The design that was merged defers the entire I/O completion, in the block layer rather than in the page cache. A bio is the block layer’s unit of I/O: a set of pages, a range on a device, and a completion callback. A submitter that knows that its callback needs task context sets a flag, BIO_COMPLETE_IN_TASK, on the bio. When the bio completes, the block layer checks the flag and the current execution context:

if (bio_flagged(bio, BIO_COMPLETE_IN_TASK) && bio_in_atomic())
	__bio_complete_in_task(bio);
else if (bio->bi_end_io)
	bio->bi_end_io(bio);

If the completion is already in a context that can sleep (i.e., non-atomic context), the callback runs immediately, with no added cost. Otherwise, the bio is pushed onto a per-CPU list and a per-CPU work item is scheduled on a workqueue.

bio_endio() calls the handler directly unless the bio is flagged BIO_COMPLETE_IN_TASK and the completion is in atomic context.
bio_endio() calls the handler directly unless the bio is flagged BIO_COMPLETE_IN_TASK and the completion is in atomic context.

A worker then drains the list and runs each callback in task context. A disk’s hardware queue steers block completions to the CPU that submitted the I/O, so using per-CPU lists avoids needing a shared lock. The only synchronization necessary is disabling interrupts on the local CPU while touching the list. The bio_in_atomic() helper checks whether the current context is not preemptible or in an RCU read section, both of which count as atomic context where sleeping is not allowed.

There is a second entry point for handlers that only discover at completion time that they need to sleep. A completion callback can call bio_complete_in_task() on its own bio; if the current context is atomic, the bio is queued and the callback returns, and the worker calls it again in task context. The plan is to use this for file system error reporting: the common path completes inline, and only the error path needs to take locks. Christoph Hellwig (the NVMe maintainer and XFS developer) described these two interfaces, taken together, as “a good building block considering the whole picture” of kernel subsystems that would benefit from this behavior.

Flagged bios are moved to a per-CPU worker, where the dropbehind folios are dropped.
Flagged bios are moved to a per-CPU worker, where the dropbehind folios are dropped.

With the block layer now handling deferral, setting BIO_COMPLETE_IN_TASK for RWF_DONTCACHE is easy. The writeback path sets it whenever it adds a dropbehind page to a write, which also allowed us to remove XFS’s special handling of dropbehind pages. Reads need nothing, since a read’s pages are dropped by the reading thread itself after the copy. The block device change is then the one line that accepts the flag. The series has been merged for Linux 7.3, which should be released in mid-October.


Reusing the infrastructure

The task-context completion infrastructure has been reported on in Linux news outlets and has since been picked up by other work: zswap writeback is aiming to use it to free cold swap cache pages on completion, F2FS plans to use it to enable DONTCACHE writes on Android, Christoph Hellwig’s XFS series uses it to verify checksums on reads, and his iomap changes use the completion-time entry point for error reporting and the iomap buffered write path. As one kernel developer said:

This BIO_COMPLETE_IN_TASK thing is really nice! I’m glad we have a proper infra now and no longer have to roll our own solutions :)


Using RWF_DONTCACHE

pwritev2(fd, &iov, 1, offset, RWF_DONTCACHE);   /* or preadv2() */
sqe->rw_flags |= RWF_DONTCACHE;                 /* io_uring */

The fio benchmark tool exposes it as --uncached=1.


Conclusion

RWF_DONTCACHE gives a database on a block device the ability to keep data in the kernel’s page cache when it benefits from it, and opt out per request for the data that does not, without the downsides of direct I/O. Supporting that functionality for block devices was a one-line change once the kernel had a way to run I/O completion work in task context. Building that in the block layer rather than in the page cache meant that the task-context completion infrastructure could be immediately reused by several other kernel subsystems.

That generality came as a benefit of working with upstream kernel maintainers. Because we upstreamed the change rather than carrying it as a private patch, the maintainers who reviewed it helped create a solution that benefits all Linux users. We plan to integrate RWF_DONTCACHE into TrieDB as soon as it’s released upstream. RWF_DONTCACHE for block devices will be in Linux 7.3 (which should be released in mid-October) for anyone running workloads on a raw block device to try.

RWF_DONTCACHE for Block Devices