Skip to main content

10.1 File Attributes, Operations & Access Methods (Sequential vs Direct)

πŸ“šModule 10: File Systems & InodesTopic 10.1⏱️18 min read
🎯High-Yield For:Computer Science Foundations β€’ Systems Engineering β€’ Storage Abstractions

πŸ’‘ Core Intuition​

Imagine a corporate law firm maintaining thousands of legal contracts in a secure physical vault:

Architecture Flow

The Legal Document Access Pipeline

Contrasting physical document storage with structured kernel access protocols

πŸ’‘ Hover or click any card for deep-dive operational details
πŸ“œThe File

The Contract Dossier

Logical File Abstraction

A continuous stack of printed legal agreements bound together under a single case name.

β†’
Check Out Document
πŸ“–The Checkout Desk

The Vault Checkout Ledger

Open-File Table

When an attorney requests the dossier, the clerk logs an active checkout token.

β†’
Reading Methods
πŸ”Reading Method

Sequential Reading vs Page Jumping

Sequential vs Direct Access

A junior associate reads the contract line-by-line from Page 1 to Page 50.

  • File: The uniform logical abstraction provided by the OS to store contiguous streams of persistent bytes across disparate storage media.
  • File Attributes: The administrative metadata (size, permissions, timestamps, owner) decoupled from file data.
  • Open-File Tables: Dual-level kernel tracking structures that maintain active read/write offsets across concurrent threads.

πŸ’» Bridging to Computer Science​

To user applications, secondary storage must not look like physical platter sectors, magnetic tracks, or flash wear-leveling blocks. The operating system provides a clean, device-independent abstraction: the File.

The 5 Hierarchical Layers of a Modern File System

From user-space system call interfaces down to physical storage hardware

Layer 5

Application Program Interface

↓
Layer 4

Logical File System

↓
Layer 3

File Organization Module

↓
Layer 2

Basic File System

↓
Layer 1

I/O Control & Device Drivers


πŸ“š Core Deep-Dive & Concepts​

1. File Attributes (Metadata)​

A file consists of two distinct components:

  1. File Data: The uninterpreted stream of bytes written by the user application.
  2. File Attributes (Metadata): Administrative information maintained by the operating system kernel.

2. Dual-Level Open-File Table Architecture​

Why doesn't the kernel perform a disk directory search every time an application calls read() or write()?

Because traversing a directory path string across disk blocks is painfully slow (5–20Β ms5\text{–}20\text{ ms}). To achieve nanosecond I/O dispatching, the OS uses Dual-Level Open-File Tables:

Dual-Level Open-File Table Architecture

Kernel decoupling of process file descriptors, dynamic file seek offsets, and persistent on-disk inodes

User Process Space

Per-Process File Descriptor Table (PCB)

↓
Kernel Global Space

System-Wide Open-File Table (struct file)

↓
Filesystem Storage Core

In-Memory Inode Table (struct inode)

  1. Per-Process File Descriptor Table:
    • Stored in the process's PCB (struct task_struct in Linux).
    • An array indexed by an integer File Descriptor (fd).
    • Contains a pointer to an entry in the System-Wide Open-File Table.
  2. System-Wide Open-File Table (Kernel Global):
    • Contains an entry for every active open instance of a file across the entire operating system.
    • Holds the Current File Offset (Seek Pointer): Tracks where the next read or write operation will take place.
    • Holds access mode flags (O_RDONLY, O_WRONLY, O_SYNC).
    • Maintains a Reference Count: Incremented when a child process inherits a file descriptor via fork().
  3. In-Memory Inode Table:
    • Caches the file's disk inode in RAM, storing file ownership, permissions, and raw physical disk block pointers.

3. File Access Methods: Sequential vs Direct (Random)​

How user programs traverse the byte stream of a file determines both filesystem design and application performance:


4. Core File System Calls (POSIX API)​

// 1. Open an existing file or create if missing
int fd = open("/var/log/syslog", O_RDWR | O_CREAT, 0644);

// 2. Reposition the seek pointer to an arbitrary byte offset (Direct Access)
off_t new_pos = lseek(fd, 4096, SEEK_SET); // Jumps straight to byte 4096

// 3. Read 512 bytes starting from current seek pointer
char buffer[512];
ssize_t bytes_read = read(fd, buffer, sizeof(buffer));

// 4. Write data to disk
ssize_t bytes_written = write(fd, buffer, bytes_read);

// 5. Commit in-flight kernel page cache buffers to physical storage media
fsync(fd);

// 6. Release file descriptor token
close(fd);

🏭 In The Real World: Production Case Study​

High-Throughput Linux I/O: read() vs pread() & io_uring​

In modern multi-threaded server applications (such as NGINX, MongoDB, or RocksDB), concurrent threads reading from the same file encounter severe contention on the seek pointer:

  1. Atomic Direct Access via pread():
    • To eliminate race conditions without acquiring mutex locks, POSIX introduced pread(fd, buf, count, offset).
    • pread() performs the seek and read in a single atomic kernel operation without updating the file descriptor's internal offset, allowing 64 threads to read concurrently from the same fd.
  2. Next-Generation Asynchronous I/O (io_uring):
    • Modern Linux replaces blocking syscalls with lockless shared-memory ring buffers (io_uring), submitting millions of direct-access I/O requests per second with zero syscall overhead.

🎯 Exam & Interview Pitfall Check​

Core Conceptual Questions

Question 1: Explain what happens when a parent process opens a file and then executes fork(). Do parent and child share the same file offset pointer? Answer:

  1. Yes! When fork() is called, the child inherits an exact duplicate of the parent's file descriptor table.
  2. Crucially, the child's file descriptor entry points to the exact same entry in the System-Wide Open-File Table in the kernel.
  3. Because the current file offset pointer resides in the System-Wide Open-File Table entry (not in the per-process table), the parent and child share the identical file offset.
  4. If the parent reads 100 bytes, the offset advances by 100. If the child then immediately calls read(), it reads from byte 100 onward.

Question 2: What is the difference between sequential access and direct access in file systems? Answer:

  • Sequential Access: Bytes or records must be read in linear order from beginning to end. To read byte 10,00010,000, the application must read all preceding 9,9999,999 bytes.
  • Direct Access: Any arbitrary byte or block can be accessed immediately by specifying its record number or byte offset (e.g. using lseek()), without reading intermediate blocks.
Common Interview Traps
  • The "File Name is Stored in the Inode" Trap: In UNIX/Linux systems, the filename is NOT stored in the inode! The inode stores file size, permissions, timestamps, and block pointers. The filename is stored exclusively inside the Directory Entry (dirent) mapping [Filename -> Inode Number].
  • Confusing Logical Size with Physical Disk Allocation: A file can have a logical size of 1Β TB1\text{ TB} (created via lseek(fd, 1TB, SEEK_SET); write(fd, "A", 1);), but on disk it occupies only a single 4Β KB4\text{ KB} block. This is a Sparse File.

πŸ’¬

Discussion & Doubts