The question: direct I/O only works if your buffer, your offset, and your length are all aligned. Aligned to what, exactly, and who decides?
The answer is per-filesystem, and the kernel will tell you if you ask correctly. Asking correctly is the whole chapter, and it is shorter than it sounds.
dioalign <path> prints a file's basic stat data plus its direct-I/O alignment
requirements, and says so plainly when the filesystem has none.
One syscall does it: statx(2)
with the STATX_DIOALIGN mask.
Chapter 4 reuses this, because you cannot write a working O_DIRECT read until
you know these two numbers.
- A file on ext4 and a file on tmpfs. Do both report an alignment?
- You pass
STATX_DIOALIGNin the request mask. Is the kernel obliged to fill those fields in? struct statx stx = {0};and the field reads back as 0. What are the two different things that could mean?
make ch0
./ch0/dioalign /mnt/work/some.datOn ext4, backed by a loop device:
$ ./ch0/dioalign /mnt/work/readme-scratch/small.dat
1786126366964 [INFO ] [dioalign] path: /mnt/work/readme-scratch/small.dat
blksize: 4096, mode: 100644, ino: 18, size: 1048576, blocks: 2048
dio_mem_align: 512, dio_offset_align: 512
The same program against a file on tmpfs:
$ ./ch0/dioalign /tmp/tiny.txt
1786126366975 [INFO ] [dioalign] path: /tmp/tiny.txt
blksize: 4096, mode: 100644, ino: 138, size: 26, blocks: 8
dio: unsupported on this filesystem (STATX_DIOALIGN not returned)
And on something that does not exist:
$ ./ch0/dioalign /mnt/work/nope
1786126366984 [ERROR] [dioalign] statx failed: No such file or directory
dio_mem_align: 512 is a requirement on the address of your buffer in memory.
dio_offset_align: 512 is a requirement on the file offset and the length of
the read. They are separate numbers because they constrain separate things, and
on some setups they differ.
blocks: 2048 counts 512-byte units, not blksize units, which is why a 1 MiB
file reports 2048 and not 256. This trips up everyone once.
The tmpfs line is the point of the chapter. tmpfs has no block device under it,
so there is no alignment to report, so the kernel silently declines to fill those
fields. It does not fail. It does not warn. It clears the bit in stx_mask and
returns success.
That is why the program checks:
if (stx.stx_mask & STATX_DIOALIGN)
printf("dio_mem_align: %u, ...", stx.stx_dio_mem_align, ...);
else
printf("dio: unsupported on this filesystem ...");Without that check, the struct was zero-initialised, so the unfilled field reads
back as 0, and 0 is indistinguishable from a real answer of zero. You would
print "alignment: 0", believe it, and spend chapter 4 debugging EINVAL from a
read whose alignment math was built on a field the kernel never wrote.
The transferable part: statx request masks are a request, not a contract.
Every field you read has to be gated on the returned stx_mask. The API is
designed to be extended, so "I asked for it" and "I got it" are two different
facts.
STATX_DIOALIGNneeds Linux 6.1 or newer. On an older kernel every file looks like tmpfs above.- A reported alignment of 0 on a real block-backed filesystem is a real answer meaning "unknown". Assume 512 and write it down. Do not go debugging your program.
#define _GNU_SOURCEmust be the first line of the file, before any include.statxand half ofstruct statxare hidden behind it. Put it after the includes and you getimplicit declaration of function 'statx'for a function whose man page is open on the other monitor.stat -freports the same magic (0xEF53) for ext2, ext3 and ext4, so it cannot tell you which filesystem you are actually on. Usemount.
man 2 statx, and specifically the RETURN VALUE section onstx_mask. That paragraph is the chapter.