# Introduce write\_at/write\_all\_at/read\_at/read\_exact\_at on Windows

**URL:** <https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649>\
**Category:** libs\
**Created:** [October 5, 2023, 8:05am UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649 "2023-10-05T08:05:01Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![nazar-pc](https://sea2.discourse-cdn.com/flex002/user_avatar/internals.rust-lang.org/nazar-pc/32/11386_2.png) [@nazar-pc](https://internals.rust-lang.org/u/nazar-pc)\
**Post date:** [October 5, 2023, 8:05am UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/1 "2023-10-05T08:05:01Z")

</div>

I have discovered that `std::os::unix::fs::FileExt::read_exact_at` and `std::os::windows::fs::FileExt::seek_read` are very far from being similar, Windows version is MUCH slower.

There was [Tracking issue for write\_all\_at/read\_exact\_at convenience methods · Issue #51984 · rust-lang/rust · GitHub](https://github.com/rust-lang/rust/issues/51984) a while ago that introduced those methods for unix, I think there would be a significant value to introduce similar/the same methods for Windows as well, especially considering that doing it in performant way seems to require unsafe and is not trivial: [https://github.com/vasi/positioned-io/blob/1cb70389b133f1e96c68279b04c6baf618e6ca22/src/windows.rs](https://github.com/vasi/positioned-io/blob/1cb70389b133f1e96c68279b04c6baf618e6ca22/src/windows.rs)

For context my use case involves reading from 1GiB sector chunks that are 32 bytes each, at random, 1MiB total. It has a pretty good performance on Linux with high read concurrency available on modern SSDs, but takes orders of magnitude more time on Windows.

---

<div class="post-metadata">

**Author:** ![zackw](https://sea2.discourse-cdn.com/flex002/user_avatar/internals.rust-lang.org/zackw/32/2071_2.png) [@zackw](https://internals.rust-lang.org/u/zackw)\
**Post date:** [October 5, 2023, 12:30pm UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/2 "2023-10-05T12:30:59Z")

</div>

`FileExt::seek_read` ultimately calls [`sys::windows::Handle::synchronous_read`](https://github.com/rust-lang/rust/blob/5c3a0e932b7c6864f98dac739b576e9ff5913739/library/std/src/sys/windows/handle.rs#L227), which calls `NtReadFile`; I'm not a Windows expert, but, as used here, this system call appears to be effectively the same thing as Unix `pread`. The code you linked as "the performant way", however, does a `MapViewOfFile` and an `UnmapViewOfFile` every time it's called. Modifying a process's memory map is inherently expensive at the hardware level, so, I would expect the technique used by `positioned-io` to be measurably _slower_ than what the stdlib is doing. You might want to look into alternative explanations for the slowdown you're observing.

Independent of that, I agree that there's no good reason for the _API_ for positioned read/write presented by `std::os::unix::fs::FileExt` to be different from that presented by `std::os::windows::fs::FileExt`. Can we converge them?

---

<div class="post-metadata">

**Author:** ![the8472](https://avatars.discourse-cdn.com/v4/letter/t/0ea827/32.png) [@the8472](https://internals.rust-lang.org/u/the8472)\
**Post date:** [October 5, 2023, 12:36pm UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/3 "2023-10-05T12:36:19Z")

</div>

windows:

> Even when the I/O Manager is maintaining the current file position, the caller can reset this position by passing an explicit _ByteOffset_ value to **NtReadFile**. Doing this automatically changes the current file position to that _ByteOffset_ value, performs the read operation, and then updates the position according to the number of bytes actually read. This technique gives the caller atomic seek-and-read service.

unix:

> ```
> pread() reads up to count bytes from file descriptor fd at offset offset (from the start of the file) into the buffer starting at buf.
> The file offset is not changed.
> 
> ```

These are just not compatible

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://sea2.discourse-cdn.com/flex002/user_avatar/internals.rust-lang.org/nazar-pc/32/11386_2.png) [@nazar-pc](https://internals.rust-lang.org/u/nazar-pc)\
**Post date:** [October 5, 2023, 12:45pm UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/4 "2023-10-05T12:45:35Z")

</div>

I must admit I have not actually tested performance of that crate, but it seemed like it should have been fast, I might be completely wrong though.

What I did test and what was much faster is memory-mapped I/O, but if mapping the whole file it was too much (we saw files that are over 60TiB in size). Currently testing mapping individual sectors of the file instead, but then we're doing file mapping a lot of times and ultimately we don't want memory-mapped I/O due to its numerous drawbacks, it seems like there must be a way to simply read the bytes I need without going through unnecessary kernel abstractions.

---

<div class="post-metadata">

**Author:** ![zackw](https://sea2.discourse-cdn.com/flex002/user_avatar/internals.rust-lang.org/zackw/32/2071_2.png) [@zackw](https://internals.rust-lang.org/u/zackw)\
**Post date:** [October 5, 2023, 1:12pm UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/5 "2023-10-05T13:12:14Z")

</div>

It might be worth experimenting with Windows' asynchronous (aka "overlapped") I/O API to queue up a whole bunch of these reads and let the kernel complete them in the most efficient order. The stdlib doesn't have anything for that but there's probably at least one crate out there.

---

<div class="post-metadata">

**Author:** ![zackw](https://sea2.discourse-cdn.com/flex002/user_avatar/internals.rust-lang.org/zackw/32/2071_2.png) [@zackw](https://internals.rust-lang.org/u/zackw)\
**Post date:** [October 5, 2023, 1:14pm UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/6 "2023-10-05T13:14:56Z")

</div>

> [@the8472](#):
>
> > **NtReadFile** [...] automatically changes the current file position to that _ByteOffset_ value,
> 
> ...
> 
> > pread() reads ... at offset offset ...The file offset is not changed.

That doesn't mean there's no way to converge the API presented by the stdlib, it just means `NtReadFile` used in this fashion doesn't provide the semantics required to match `pread`. I'm curious what the alternative to "Even when the I/O Manager is maintaining the current file position" is.

---

<div class="post-metadata">

**Author:** ![the8472](https://avatars.discourse-cdn.com/v4/letter/t/0ea827/32.png) [@the8472](https://internals.rust-lang.org/u/the8472)\
**Post date:** [October 5, 2023, 2:07pm UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/7 "2023-10-05T14:07:29Z")

</div>

Another option would be [io\_uring](https://man7.org/linux/man-pages/man7/io_uring.7.html) or its [windows equivalent](https://learn.microsoft.com/en-us/windows/win32/api/ioringapi/) which should minimize the number of syscalls for those small reads.

> [@nazar-pc](#):
>
> For context my use case involves reading from 1GiB sector chunks that are 32 bytes each, at random, 1MiB total

Yeah, 32k randomly scattered 32byte reads are pretty bad. I would expect even `pread` to be suboptimal here just due to the syscall overhead.

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://sea2.discourse-cdn.com/flex002/user_avatar/internals.rust-lang.org/nazar-pc/32/11386_2.png) [@nazar-pc](https://internals.rust-lang.org/u/nazar-pc)\
**Post date:** [October 5, 2023, 2:48pm UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/8 "2023-10-05T14:48:50Z")

</div>

Yeah, `io_uring` is the ultimate goal, but it is only on Linux and requires fairly new kernel (many users are on Ubuntu 20.04 with non-hwe kernel unfortunately).

There was an idea to just read the whole gigabyte and drop unnecessary data, which would be actually faster on Windows that what I have experienced with `seek_read`, but it wouldn't be fast enough on SATA SSDs and we have tight time requirements for this operation.

Modern SSDs are shockingly good at serving multiple concurrent requests, just need to find a way to leverage it on 3 major platforms (yes, macOS is also a target).

---

<div class="post-metadata">

**Author:** ![chrisd](https://sea2.discourse-cdn.com/flex002/user_avatar/internals.rust-lang.org/chrisd/32/7232_2.png) [@chrisd](https://internals.rust-lang.org/u/chrisd)\
**Post date:** [October 5, 2023, 3:06pm UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/9 "2023-10-05T15:06:35Z")

</div>

> [@zackw](#):
>
> I'm curious what the alternative to "Even when the I/O Manager is maintaining the current file position" is.

Files that aren't seekable don't maintain a file pointer. Files opened for async (overlapped) access do not need to maintain a file position.

---

<div class="post-metadata">

**Author:** ![chrisd](https://sea2.discourse-cdn.com/flex002/user_avatar/internals.rust-lang.org/chrisd/32/7232_2.png) [@chrisd](https://internals.rust-lang.org/u/chrisd)\
**Post date:** [October 5, 2023, 3:16pm UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/10 "2023-10-05T15:16:12Z")

</div>

I would strongly suggest coming up with some benchmarks first. Improving performance by guesswork is a lot more tricky.

That said, I doubt it'd be hard to beat std for read/write perf. All the std APIs are designed to be synchronous and to work on any open handle, so using IOCP would be a great improvement. Even better would be IoRing but that requires a newer kernel.

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://sea2.discourse-cdn.com/flex002/user_avatar/internals.rust-lang.org/nazar-pc/32/11386_2.png) [@nazar-pc](https://internals.rust-lang.org/u/nazar-pc)\
**Post date:** [October 20, 2023, 10:16am UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/11 "2023-10-20T10:16:31Z")

</div>

Implemented benchmark (not reduced for now, just in my app) and ran with both memory-maped I/O and `seek_read` on Windows. The result is that memory-mapped I/O results in 0% idle time on a disk, with `seek_read` idle time is ~15% and total execution time of the app is much longer: 71ms vs 267ms.

Is anyone available of the non-async library for Windows for reading files at arbitrary offset without seeking first in the meantime (on Linux async version of the app is slower anyway)?

---

<div class="post-metadata">

**Author:** ![the8472](https://avatars.discourse-cdn.com/v4/letter/t/0ea827/32.png) [@the8472](https://internals.rust-lang.org/u/the8472)\
**Post date:** [October 20, 2023, 3:04pm UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/12 "2023-10-20T15:04:54Z")

</div>

If you're doing synchronous IO you'll either want to issue readaheads or throw more threads at the problem so that the parallelism can mask the wait-induced stalls

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://sea2.discourse-cdn.com/flex002/user_avatar/internals.rust-lang.org/nazar-pc/32/11386_2.png) [@nazar-pc](https://internals.rust-lang.org/u/nazar-pc)\
**Post date:** [October 20, 2023, 3:48pm UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/13 "2023-10-20T15:48:36Z")

</div>

Reads are random, number of threads is already equal to number of cores and works great on Linux and macOS. My current suspicion is that `seek_read` has the "seek" part, so calling it concurrently from multiple threads can be problematic, I'll try to open file multiple times, once in each thread, on Windows and see how that goes.

---

<div class="post-metadata">

**Author:** ![zackw](https://sea2.discourse-cdn.com/flex002/user_avatar/internals.rust-lang.org/zackw/32/2071_2.png) [@zackw](https://internals.rust-lang.org/u/zackw)\
**Post date:** [October 20, 2023, 7:25pm UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/14 "2023-10-20T19:25:31Z")

</div>

The memory-mapped IO version, are you making one big map of the entire file and keeping it around for the duration of whatever it is your app actually does, or are you doing what `positioned-io` does, mapping and unmapping chunks as needed?

---

<div class="post-metadata">

**Author:** ![chrisd](https://sea2.discourse-cdn.com/flex002/user_avatar/internals.rust-lang.org/chrisd/32/7232_2.png) [@chrisd](https://internals.rust-lang.org/u/chrisd)\
**Post date:** [October 20, 2023, 10:04pm UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/15 "2023-10-20T22:04:00Z")

</div>

Memory mapped I/O is great! However, being realistic it's not something std can do by default behind the user's back. Crates have a lot more freedom though!

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://sea2.discourse-cdn.com/flex002/user_avatar/internals.rust-lang.org/nazar-pc/32/11386_2.png) [@nazar-pc](https://internals.rust-lang.org/u/nazar-pc)\
**Post date:** [October 23, 2023, 7:49am UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/16 "2023-10-23T07:49:25Z")

</div>

> [@zackw](#):
>
> The memory-mapped IO version, are you making one big map of the entire file and keeping it around for the duration of whatever it is your app actually does, or are you doing what `positioned-io` does, mapping and unmapping chunks as needed?

Mapping the whole file for the duration of the app lifetime otherwise performance is even worse than direct file reads. The challenge is that files are multiple (sometimes tens) of terabytes in size, so Windows fill memory with pages (and users are concerned with application using 100% of memory available even though it is not quite the case), also there is a limit of how much space can be mapped this way in total apparently that seems to be smaller than amout of virtual memory that some users are hitting as well (though it might be Linux-specific, I'm not sure).

> [@chrisd](#):
>
> Memory mapped I/O is great! However, being realistic it's not something std can do by default behind the user's back. Crates have a lot more freedom though!

I'm not suggesting it either, I'm just sharing what seems to work better and that there are APIs on other platforms in `std` that do exactly what I need without memory-mapped I/O efficiently.

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://sea2.discourse-cdn.com/flex002/user_avatar/internals.rust-lang.org/nazar-pc/32/11386_2.png) [@nazar-pc](https://internals.rust-lang.org/u/nazar-pc)\
**Post date:** [October 23, 2023, 5:46pm UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/17 "2023-10-23T17:46:34Z")

</div>

Opening file multiple times, once for every thread in a thread pool works even better than memory-mapped I/O, somewhat expectedly. I think I'll go with that now. Seems to be tiny bit faster on Linux with `read_exact_at` as well, but hardly outside of noise range.

On Windows relative numbers are:

- 936 ms for single file with `seek_read` (+handling of partial reads identically to `read_exact_at` in std)
- 245 ms for single file with memory-mapped I/O
- 190 ms with file openes many times, once for each thead in thread pool using `seek_read`

This was on Micron 5200 3.8T SATA SSD and i7-6700 processor.

---

<div class="post-metadata">

**Author:** ![zackw](https://sea2.discourse-cdn.com/flex002/user_avatar/internals.rust-lang.org/zackw/32/2071_2.png) [@zackw](https://internals.rust-lang.org/u/zackw)\
**Post date:** [October 23, 2023, 5:53pm UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/18 "2023-10-23T17:53:27Z")

</div>

Can you post your test program somewhere? The numbers for memory-mapped I/O with a map and unmap for each read are so far off from what I'd expect, that I want to run the same benchmark on a couple different Unixes to find out if my mental model of the cost of altering a process's address space is really that wrong.

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://sea2.discourse-cdn.com/flex002/user_avatar/internals.rust-lang.org/nazar-pc/32/11386_2.png) [@nazar-pc](https://internals.rust-lang.org/u/nazar-pc)\
**Post date:** [October 23, 2023, 6:08pm UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/19 "2023-10-23T18:08:44Z")

</div>

I'm afraid the program will not be very useful as it requires long preparation step and then audits results of preparation.

Essentially I have [this trait](https://github.com/subspace/subspace/blob/c14b3ff0d6a2547f4d9155f1acf231f3181dc68d/crates/subspace-farmer-components/src/lib.rs#L67-L84):

> ****
>
> ```rust
> pub trait ReadAtSync: Send + Sync {
> /// Fill the buffer by reading bytes at a specific offset
> fn read_at(&self, buf: &mut [u8], offset: usize) -> io::Result<()>;
> }
> 
> ```

It is implemented for [`&[u8]`](https://github.com/subspace/subspace/blob/c14b3ff0d6a2547f4d9155f1acf231f3181dc68d/crates/subspace-farmer-components/src/lib.rs#L160-L173):

> ****
>
> ```rust
> impl ReadAtSync for [u8] {
> fn read_at(&self, buf: &mut [u8], offset: usize) -> io::Result<()> {
> if buf.len() + offset > self.len() {
> return Err(io::Error::new(
> io::ErrorKind::InvalidInput,
> "Buffer length with offset exceeds own length",
> ));
> }
> 
> buf.copy_from_slice(&self[offset..][..buf.len()]);
> 
> Ok(())
> }
> }
> 
> ```

And [`File`](https://github.com/subspace/subspace/blob/c14b3ff0d6a2547f4d9155f1acf231f3181dc68d/crates/subspace-farmer-components/src/lib.rs#L202-L206):

> ****
>
> ```rust
> impl ReadAtSync for File {
> fn read_at(&self, buf: &mut [u8], offset: usize) -> io::Result<()> {
> self.read_exact_at(buf, offset as u64)
> }
> }
> 
> ```

And then I have created a [following wrapper](https://github.com/subspace/subspace/blob/3f5e7dccb511325cb12361b0441fb4a33ac7f3e2/crates/subspace-farmer/src/single_disk_farm/farming/rayon_files.rs) that opens file multiple times and implements the same trait leveraging above `File` implementation of the trait:

> ****
>
> ```rust
> 
> pub struct RayonFiles {
> files: Vec<File>,
> }
> 
> impl ReadAtSync for RayonFiles {
> fn read_at(&self, buf: &mut [u8], offset: usize) -> io::Result<()> {
> let thread_index = rayon::current_thread_index().ok_or_else(|| {
> io::Error::new(
> io::ErrorKind::Other,
> "Reads must be called from rayon worker thread",
> )
> })?;
> let file = self.files.get(thread_index).ok_or_else(|| {
> io::Error::new(io::ErrorKind::Other, "No files entry for this rayon thread")
> })?;
> 
> file.read_at(buf, offset)
> }
> }
> 
> impl RayonFiles {
> pub fn open(path: &Path) -> io::Result<Self> {
> let files = (0..rayon::current_num_threads())
> .map(|_| {
> let file = OpenOptions::new()
> .read(true)
> .advise_random_access()
> .open(path)?;
> file.advise_random_access()?;
> 
> Ok::<_, io::Error>(file)
> })
> .collect::<Result<Vec<_>, _>>()?;
> 
> Ok(Self { files })
> }
> }
> 
> ```

All file operations above are proving OS hints about random file reads using these cross-platform utility traits (they implement cross-platform version of `read_exact_at` as described before): [https://github.com/subspace/subspace/blob/c14b3ff0d6a2547f4d9155f1acf231f3181dc68d/crates/subspace-farmer-components/src/file\_ext.rs](https://github.com/subspace/subspace/blob/c14b3ff0d6a2547f4d9155f1acf231f3181dc68d/crates/subspace-farmer-components/src/file_ext.rs)

Now what I actually do with that is running some CPU-intensive work interleaved with random file reads (~20kiB each per 1GB of space used) using rayon on a large file (2.7TB in above case).

So above numbers are not just disk reads, they represent the workload I actually care about where reads are only a component, but it is clear that there is a massive difference depending on how files are read on Windows, it is even larger than above results show due to CPU being a significant contributor to the performance.

---

<div class="post-metadata">

**Author:** ![the8472](https://avatars.discourse-cdn.com/v4/letter/t/0ea827/32.png) [@the8472](https://internals.rust-lang.org/u/the8472)\
**Post date:** [October 23, 2023, 9:54pm UTC](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649/20 "2023-10-23T21:54:51Z")

</div>

> [@zackw](#):
>
> find out if my mental model of the cost of altering a process's address space is really that wrong.

Linux only gained finer-grained locking fairly recently. And I think there also have been some optimizations to reduce IPIs/TLB shootdowns. So it can depend on which kernel version you're using.

> **[Concurrent page-fault handling with per-VMA locks \[LWN.net\]](https://lwn.net/Articles/906852/)**
>
> The kernel is, in many ways, a marvel of scalability, but there is a
> longstanding pain point in the memory-management subsystem that has
> resisted all attempts at elimination: the mmap\_lock. This lock
> was inevitably a topic at the 2022 Linux
> Storage,...

```rust
$ zcat /proc/config.gz | grep CONFIG_PER_VMA_LOCK
CONFIG_PER_VMA_LOCK=y

```

[Next page](https://internals.rust-lang.org/t/introduce-write-at-write-all-at-read-at-read-exact-at-on-windows/19649.md?page=2)
