core/sys/linux/uring
uring
Types
3Completion_Queue
Completion_Queue :: struct {
head: ^u32,
tail: ^u32,
mask: u32,
overflow: ^u32,
cqes: []linux.IO_Uring_CQE,
}SourceRing
Ring :: struct {
fd: linux.Fd,
sq: Submission_Queue,
cq: Completion_Queue,
flags: linux.IO_Uring_Setup_Flags,
features: linux.IO_Uring_Features,
}SourceSubmission_Queue
Submission_Queue :: struct {
head: ^u32,
tail: ^u32,
mask: u32,
flags: ^linux.IO_Uring_Submission_Queue_Flags,
dropped: ^u32,
array: []u32,
sqes: []linux.IO_Uring_SQE,
mmap: []u8,
mmap_sqes: []u8,
// We use `sqe_head` and `sqe_tail` in the same way as liburing:
// We increment `sqe_tail` (but not `tail`) for each call to `get_sqe()`.
// We then set `tail` to `sqe_tail` once, only when these events are actually submitted.
// This allows us to amortize the cost of the @atomicStore to `tail` across multiple SQEs.
sqe_head: u32,
sqe_tail: u32,
}SourceConstants
4DEFAULT_ENTRIES
DEFAULT_ENTRIES :: 32SourceDEFAULT_PARAMS
DEFAULT_PARAMS :: linux.IO_Uring_Params = linux.IO_Uring_Params {
sq_thread_idle = DEFAULT_THREAD_IDLE_MS,
}SourceDEFAULT_THREAD_IDLE_MS
DEFAULT_THREAD_IDLE_MS :: 1000SourceMAX_ENTRIES
MAX_ENTRIES :: 4096SourceProcedures
79accept
accept :: proc(
ring: ^Ring,
user_data: u64,
sockfd: linux.Fd,
addr: ^T,
addr_len: ^i32,
flags: linux.Socket_FD_Flags,
file_index: u32,
) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceIssue the equivalent of an accept4(2) system call.
See also accept4(2) for the general description of the related system call.
If the file_index field is set to a positive number, the file won't be installed into the normal file table as usual but will be placed into the fixed file table at index file_index - 1. In this case, instead of returning a file descriptor, the result will contain either 0 on success or an error. If the index points to a valid empty slot, the installation is guaranteed to not fail. If there is already a file in the slot, it will be replaced, similar to IORING_OP_FILES_UPDATE. Please note that only uring has access to such files and no other syscall can use them. See IOSQE_FIXED_FILE and IORING_REGISTER_FILES.
Available since 5.5.
async_cancel
async_cancel :: proc(ring: ^Ring, orig_user_data: u64, user_data: u64) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceAttempt to cancel an already issued request.
The request is identified by it's user data.
The cancelation request will complete with one of the following results codes.
If found, the res field of the cqe will contain 0. If not found, res will contain -ENOENT.
If found and attempted canceled, the res field will contain -EALREADY. In this case, the request may or may not terminate. In general, requests that are interruptible (like socket IO) will get canceled, while disk IO requests cannot be canceled if already started.
Available since 5.5.
bind
bind :: proc(ring: ^Ring, user_data: u64, sock: linux.Fd, addr: ^T) -> (sqe: linux.IO_Uring_SQE, ok: bool)SourceIssues the equivalent of the bind(2) system call.
Available since 6.11.
close
close :: proc(ring: ^Ring, user_data: u64, fd: linux.Fd, file_index: u32) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceIssue the equivalent of a close(2) system call.
See also close(2) for the general description of the related system call.
Available since 5.6.
If the file_index field is set to a positive number, this command can be used to close files that were direct opened through IORING_OP_OPENAT, IORING_OP_OPENAT2, or IORING_OP_ACCEPT using the uring specific direct descriptors. Note that only one of the descriptor fields may be set. The direct close feature is available since the 5.15 kernel, where direct descriptors were introduced.
completion_queue_make
completion_queue_make :: proc(fd: linux.Fd, params: ^linux.IO_Uring_Params, sq: ^Submission_Queue) -> (Completion_Queue)Sourceconnect
connect :: proc(ring: ^Ring, user_data: u64, sockfd: linux.Fd, addr: ^T) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceIssue the equivalent of a connect(2) system call.
See also connect(2) for the general description of the related system call.
Available since 5.5.
copy_cqes
copy_cqes :: proc(ring: ^Ring, cqes: []linux.IO_Uring_CQE, wait_nr: u32) -> (n_copied: u32, err: linux.Errno)SourceCopies as many CQEs as are ready, and that can fit into the destination cqes slice. If none are available, enters into the kernel to wait for at most wait_nr CQEs. Returns the number of CQEs copied, advancing the CQ ring. Provides all the wait/peek methods found in liburing, but with batching and a single method. TODO: allow for timeout.
copy_cqes_ready
copy_cqes_ready :: proc(ring: ^Ring, cqes: []linux.IO_Uring_CQE) -> (n_copied: u32)Sourcecq_advance
cq_advance :: proc(ring: ^Ring, count: u32)SourceFor advanced use cases only that implement custom completion queue methods. Matches the implementation of cq_advance() in liburing.
cq_ready
cq_ready :: proc(ring: ^Ring) -> (n_ready: u32)SourceReturns the number of completion queue entries in the completion queue (yet to consume).
cq_ring_needs_flush
cq_ring_needs_flush :: proc(ring: ^Ring) -> (bool)Sourcecqe_seen
cqe_seen :: proc(ring: ^Ring)SourceFor advanced use cases only that implement custom completion queue methods. If you use copy_cqes() or copy_cqe() you must not call cqe_seen() or cq_advance(). Must be called exactly once after a zero-copy CQE has been processed by your application. Not idempotent, calling more than once will result in other CQEs being lost. Matches the implementation of cqe_seen() in liburing.
destroy
destroy :: proc(ring: ^Ring)Sourceepoll_ctl
epoll_ctl :: proc(
ring: ^Ring,
user_data: u64,
epfd: linux.Fd,
op: linux.EPoll_Ctl_Opcode,
fd: linux.Fd,
event: ^linux.EPoll_Event,
) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceAdd, remove or modify entries in the interest list of epoll(7).
See epoll_ctl(2) for details of the system call.
Available since 5.6.
fadvise
fadvise :: proc()Sourcefallocate
fallocate :: proc()Sourcefgetxattr
fgetxattr :: proc()Sourcefiles_update
files_update :: proc(ring: ^Ring, user_data: u64, fds: []linux.Fd, off: u64) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceThis command is an alternative to using IORING_REGISTER_FILES_UPDATE which then works in an async fashion, like the rest of the uring commands.
Note that the array of file descriptors pointed to in addr must remain valid until this operation has completed.
Available since 5.6.
fixed_fd_install
fixed_fd_install :: proc()Sourcefixed_file
fixed_file :: proc()Sourceflush_sq
flush_sq :: proc(ring: ^Ring) -> (n_pending: u32)SourceSync internal state with kernel ring state on the submission queue side. Returns the number of all pending events in the submission queue. Rationale is to determine that an enter call is needed.
free_space
free_space :: proc(ring: ^Ring) -> (int)Sourcefsetxattr
fsetxattr :: proc()Sourcefsync
fsync :: proc(ring: ^Ring, user_data: u64, fd: linux.Fd, flags: linux.IO_Uring_Fsync_Flags) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceFile sync. See also fsync(2).
Optionally off and len can be used to specify a range within the file to be synced rather than syncing the entire file, which is the default behavior.
Note that, while I/O is initiated in the order in which it appears in the submission queue, completions are unordered. For example, an application which places a write I/O followed by an fsync in the submission queue cannot expect the fsync to apply to the write. The two operations execute in parallel, so the fsync may complete before the write is issued to the storage. The same is also true for previously issued writes that have not completed prior to the fsync. To enforce ordering one may utilize linked SQEs, IOSQE_IO_DRAIN or wait for the arrival of CQEs of requests which have to be ordered before a given request before submitting its SQE.
ftruncate
ftruncate :: proc()Sourcefutex_wait
futex_wait :: proc()Sourcefutex_waitv
futex_waitv :: proc()Sourcefutex_wake
futex_wake :: proc()Sourceget_sqe
get_sqe :: proc(ring: ^Ring, extra: int) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceReturns a pointer to a vacant submission queue entry, or nil if the submission queue is full. NOTE: extra is so you can make sure there is space for related entries, defaults to 1 so a link timeout op can always be added after another.
getxattr
getxattr :: proc()Sourceinit
init :: proc(ring: ^Ring, params: ^linux.IO_Uring_Params, entries: u32) -> (err: linux.Errno)SourceInitialize and setup an uring, entries must be a power of 2 between 1 and 4096.
link_timeout
link_timeout :: proc(ring: ^Ring, user_data: u64, ts: ^linux.Time_Spec, flags: linux.IO_Uring_Timeout_Flags) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceAdds a link timeout operation.
You need to set linux.IOSQE_IO_LINK to flags of the target operation and then call this method right after the target operation. See https://lwn.net/Articles/803932/ for detail.
If the dependent request finishes before the linked timeout, the timeout is canceled. If the timeout finishes before the dependent request, the dependent request will be canceled.
The completion event result of the link_timeout will be -ETIME if the timeout finishes before the dependent request (in this case, the completion event result of the dependent request will be -ECANCELED), or -EALREADY if the dependent request finishes before the linked timeout.
Available since 5.5.
linkat
linkat :: proc()Sourcelisten
listen :: proc(ring: ^Ring, user_data: u64, fd: linux.Fd, backlog: u64) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceIssues the equivalent of the listen(2) system call.
fd must contain the file descriptor of the socket and addr must contain the backlog parameter, i.e. the maximum amount of pending queued connections.
Available since 6.11.
madvise
madvise :: proc(ring: ^Ring, user_data: u64, addr: rawptr, size: u32, advise: linux.MAdvice) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceIssue the equivalent of a madvise(2) system call.
See also madvise(2) for the general description of the related system call.
Available since 5.6.
mkdirat
mkdirat :: proc()Sourcemsg_ring
msg_ring :: proc()Sourcenop
nop :: proc(ring: ^Ring, user_data: u64) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceDo not perform any I/O. This is useful for testing the performance of the uring implementation itself.
openat
openat :: proc(
ring: ^Ring,
user_data: u64,
dirfd: linux.Fd,
path: cstring,
mode: linux.Mode,
flags: linux.Open_Flags,
file_index: u32,
) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceIssue the equivalent of a openat(2) system call.
See also openat(2) for the general description of the related system call.
Available since 5.6.
If the file_index is set to a positive number, the file won't be installed into the normal file table as usual but will be placed into the fixed file table at index file_index - 1. In this case, instead of returning a file descriptor, the result will contain either 0 on success or an error. If the index points to a valid empty slot, the installation is guaranteed to not fail. If there is already a file in the slot, it will be replaced, similar to IORING_OP_FILES_UPDATE. Please note that only uring has access to such files and no other syscall can use them. See IOSQE_FIXED_FILE and IORING_REGISTER_FILES.
Available since 5.15.
openat2
openat2 :: proc()Sourcepoll_add
poll_add :: proc(ring: ^Ring, user_data: u64, fd: linux.Fd, events: linux.Fd_Poll_Events, flags: linux.IO_Uring_Poll_Add_Flags) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourcePoll the fd specified in the submission queue entry for the events specified in the poll_events field.
Unlike poll or epoll without EPOLLONESHOT, by default this interface always works in one shot mode. That is, once the poll operation is completed, it will have to be resubmitted.
If IORING_POLL_ADD_MULTI is set in the SQE len field, then the poll will work in multi shot mode instead. That means it'll repatedly trigger when the requested event becomes true, and hence multiple CQEs can be generated from this single SQE. The CQE flags field will have IORING_CQE_F_MORE set on completion if the application should expect further CQE entries from the original request. If this flag isn't set on completion, then the poll request has been terminated and no further events will be generated. This mode is available since 5.13.
This command works like an async poll(2) and the completion event result is the returned mask of events.
Without IORING_POLL_ADD_MULTI and the initial poll operation with IORING_POLL_ADD_MULTI the operation is level triggered, i.e. if there is data ready or events pending etc. at the time of submission a corresponding CQE will be posted. Potential further completions beyond the first caused by a IORING_POLL_ADD_MULTI are edge triggered.
poll_remove
poll_remove :: proc(ring: ^Ring, user_data: u64, fd: linux.Fd, events: linux.Fd_Poll_Events) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceRemove an existing poll request.
If found, the res field of the struct io_uring_cqe will contain 0. If not found, res will contain -ENOENT, or -EALREADY if the poll request was in the process of completing already.
poll_update_events
poll_update_events :: proc(ring: ^Ring, user_data: u64, orig_user_data: u64, fd: linux.Fd, events: linux.Fd_Poll_Events) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceUpdate the events of an existing poll request.
The request will update an existing poll request with the mask of events passed in with this request. The lookup is based on the user_data field of the original SQE submitted.
Updating an existing poll is available since 5.13.
poll_update_user_data
poll_update_user_data :: proc(ring: ^Ring, user_data: u64, orig_user_data: u64, new_user_data: u64, fd: linux.Fd) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceUpdate the user data of an existing poll request.
The request will update the user_data of an existing poll request based on the value passed.
Updating an existing poll is available since 5.13.
provide_buffers
provide_buffers :: proc()Sourceread
read :: proc(ring: ^Ring, user_data: u64, fd: linux.Fd, buf: []u8, offset: u64) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceIssue the equivalent of a pread(2) system call.
If offset is set to -1 , the offset will use (and advance) the file position, like the read(2) system calls. These are non-vectored versions of the IORING_OP_READV and IORING_OP_WRITEV opcodes. See also read(2) for the general description of the related system call.
Available since 5.6.
read_fixed
read_fixed :: proc()Sourceread_multishot
read_multishot :: proc()Sourcereadv
readv :: proc(ring: ^Ring, user_data: u64, fd: linux.Fd, iovs: []linux.IO_Vec, off: u64) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceVectored read operation, see also readv(2).
recv
recv :: proc(
ring: ^Ring,
user_data: u64,
sockfd: linux.Fd,
buf: []u8,
flags: linux.Socket_Msg,
poll_first: untyped boolean = false,
) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceWorks just like send, but receives instead of sends.
poll_first: If set, uring will assume the socket is currently empty and attempting to receive data will be unsuccessful. For this case, uring will arm internal poll and trigger a receive of the data when the socket has data to be read. This initial receive attempt can be wasteful for the case where the socket is expected to be empty, setting this flag will bypass the initial receive attempt and go straight to arming poll. If poll does indicate that data is ready to be received, the operation will proceed.
Available since 5.6.
recvmsg
recvmsg :: proc(
ring: ^Ring,
user_data: u64,
fd: linux.Fd,
msghdr: ^linux.Msg_Hdr,
flags: linux.Socket_Msg,
poll_first: untyped boolean = false,
) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceWorks just like sendmsg, but receives instead of sends.
poll_first: If set, uring will assume the socket is currently empty and attempting to receive data will be unsuccessful. For this case, uring will arm internal poll and trigger a receive of the data when the socket has data to be read. This initial receive attempt can be wasteful for the case where the socket is expected to be empty, setting this flag will bypass the initial receive attempt and go straight to arming poll. If poll does indicate that data is ready to be received, the operation will proceed.
Available since 5.3.
remove_buffers
remove_buffers :: proc()Sourcerenameat
renameat :: proc()Sourcesend
send :: proc(
ring: ^Ring,
user_data: u64,
sockfd: linux.Fd,
buf: []u8,
flags: linux.Socket_Msg,
poll_first: untyped boolean = false,
) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceIssue the equivalent of a send(2) system call.
See also send(2) for the general description of the related system call.
poll_first: If set, uring will assume the socket is currently full and attempting to send data will be unsuccessful. For this case, uring will arm internal poll and trigger a send of the data when there is enough space available. This initial send attempt can be wasteful for the case where the socket is expected to be full, setting this flag will bypass the initial send attempt and go straight to arming poll. If poll does indicate that data can be sent, the operation will proceed.
Available since 5.6.
send_zc
send_zc :: proc()Sourcesendmsg
sendmsg :: proc(
ring: ^Ring,
user_data: u64,
fd: linux.Fd,
msghdr: ^linux.Msg_Hdr,
flags: linux.Socket_Msg,
poll_first: untyped boolean = false,
) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceIssue the equivalent of a sendmsg(2) system call.
See also sendmsg(2) for the general description of the related system call.
poll_first: if set, uring will assume the socket is currently full and attempting to send data will be unsuccessful. For this case, uring will arm internal poll and trigger a send of the data when there is enough space available. This initial send attempt can be wasteful for the case where the socket is expected to be full, setting this flag will bypass the initial send attempt and go straight to arming poll. If poll does indicate that data can be sent, the operation will proceed.
Available since 5.3.
sendmsg_zc
sendmsg_zc :: proc()Sourcesendto
sendto :: proc(
ring: ^Ring,
user_data: u64,
sockfd: linux.Fd,
buf: []u8,
flags: linux.Socket_Msg,
dest: ^T,
poll_first: untyped boolean = false,
) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)Sourcesetxattr
setxattr :: proc()Sourceshutdown
shutdown :: proc(ring: ^Ring, user_data: u64, fd: linux.Fd, how: linux.Shutdown_How) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceIssue the equivalent of a shutdown(2) system call.
Available since 5.11.
socket
socket :: proc(
ring: ^Ring,
user_data: u64,
domain: linux.Address_Family,
socktype: linux.Socket_Type,
protocol: linux.Protocol,
file_index: u32,
) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceIssue the equivalent of a socket(2) system call.
See also socket(2) for the general description of the related system call.
Available since 5.19.
If the file_index field is set to a positive number, the file won't be installed into the normal file table as usual but will be placed into the fixed file table at index file_index - 1. In this case, instead of returning a file descriptor, the result will contain either 0 on success or an error. If the index points to a valid empty slot, the installation is guaranteed to not fail. If there is already a file in the slot, it will be replaced, similar to IORING_OP_FILES_UPDATE. Please note that only uring has access to such files and no other syscall can use them. See IOSQE_FIXED_FILE and IORING_REGISTER_FILES.
splice
splice :: proc(
ring: ^Ring,
user_data: u64,
fd_in: linux.Fd,
off_in: i64,
fd_out: linux.Fd,
off_out: i64,
len: u32,
flags: linux.IO_Uring_Splice_Flags,
) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceIssue the equivalent of a splice(2) system call.
A sentinel value of -1 is used to pass the equivalent of a NULL for the offsets to splice(2).
Please note that one of the file descriptors must refer to a pipe. See also splice(2) for the general description of the related system call.
Available since 5.7.
sq_ready
sq_ready :: proc(ring: ^Ring) -> (u32)SourceReturns the number of submission queue entries in the submission queue.
sq_ring_needs_enter
sq_ring_needs_enter :: proc(ring: ^Ring, flags: ^linux.IO_Uring_Enter_Flags) -> (bool)SourceReturns true if we are not using an SQ thread (thus nobody submits but us), or if IORING_SQ_NEED_WAKEUP is set and the SQ thread must be explicitly awakened. For the latter case, we set the SQ thread wakeup flag. Matches the implementation of sq_ring_needs_enter() in liburing.
statx
statx :: proc(
ring: ^Ring,
user_data: u64,
dirfd: linux.Fd,
pathname: cstring,
flags: linux.FD_Flags,
mask: linux.Statx_Mask,
buf: ^linux.Statx,
) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceIssue the equivalent of a statx(2) system call.
See also statx(2) for the general description of the related system call.
Available since 5.6.
submission_queue_destroy
submission_queue_destroy :: proc(sq: ^Submission_Queue) -> (err: linux.Errno)Sourcesubmission_queue_make
submission_queue_make :: proc(fd: linux.Fd, params: ^linux.IO_Uring_Params) -> (sq: Submission_Queue, err: linux.Errno)Sourcesubmit
submit :: proc(ring: ^Ring, wait_nr: u32, timeout: ^linux.Time_Spec) -> (n_submitted: u32, err: linux.Errno)SourceSubmits the submission queue entries acquired via get_sqe(). Returns the number of entries submitted. Optionally wait for a number of events by setting wait_nr, and/or set a maximum wait time by setting timeout.
symlinkat
symlinkat :: proc()Sourcesync_file_range
sync_file_range :: proc()Sourcetee
tee :: proc(
ring: ^Ring,
user_data: u64,
fd_in: linux.Fd,
fd_out: linux.Fd,
len: u32,
flags: linux.IO_Uring_Splice_Flags,
) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceIssue the equivalent of a tee(2) system call.
Please note that both of the file descriptors must refer to a pipe. See also tee(2) for the general description of the related system call.
Available since 5.8.
timeout
timeout :: proc(ring: ^Ring, user_data: u64, ts: ^linux.Time_Spec, count: u32, flags: linux.IO_Uring_Timeout_Flags) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceRegister a timeout operation.
The timeout will complete when either the timeout expires, or after the specified number of events complete (if count is greater than 0).
flags may be 0 for a relative timeout, or IORING_TIMEOUT_ABS for an absolute timeout.
The completion event result will be -ETIME if the timeout completed through expiration, 0 if the timeout completed after the specified number of events, or -ECANCELED if the timeout was removed before it expired.
uring timeouts use the CLOCK.MONOTONIC clock source.
timeout_remove
timeout_remove :: proc(ring: ^Ring, user_data: u64, timeout_user_data: u64, flags: linux.IO_Uring_Timeout_Flags) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceRmove an existing timeout operation.
The timeout is identified by it's user_data.
The completion event result will be 0 if the timeout was found and cancelled successfully, -EBUSY if the timeout was found but expiration was already in progress, or -ENOENT if the timeout was not found.
unlinkat
unlinkat :: proc()Sourceuring_cmd
uring_cmd :: proc()Sourcewaitid
waitid :: proc()Sourcewrite
write :: proc(ring: ^Ring, user_data: u64, fd: linux.Fd, buf: []u8, offset: u64) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceIssue the equivalent of a pwrite(2) system call.
If offset is set to -1 , the offset will use (and advance) the file position, like the read(2) system calls. These are non-vectored versions of the IORING_OP_READV and IORING_OP_WRITEV opcodes. See also write(2) for the general description of the related system call.
Available since 5.6.
write_fixed
write_fixed :: proc()Sourcewritev
writev :: proc(ring: ^Ring, user_data: u64, fd: linux.Fd, iovs: []linux.IO_Vec, off: u64) -> (sqe: ^linux.IO_Uring_SQE, ok: bool)SourceVectored write operation, see also writev(2).