Skip to content

file handling

File handling covers everything your program does with a file on disk: building the path, opening it in the right mode, reading or writing its contents, and closing it again. It’s one of the first skills you pick up in Python, and one of the easiest to get subtly wrong, because most of the mistakes only show up on someone else’s machine or on a file bigger than the one you tested with.

The closing part is already solved. Open every file inside a with statement and the context manager closes it for you, even when an exception interrupts the block. That much belongs to resource management. The practices below cover the decisions that with can’t make for you.

When you’re reading and writing files in Python, you can follow the best practices below:

  • Build paths with pathlib, not string operations. Gluing fragments together with + or a hardcoded / breaks as soon as your code runs on another operating system. The pathlib module lets you write Path("logs") / "app.log" instead, and the resulting object carries methods like .read_text(), .glob(), and .mkdir().
  • Always pass encoding="utf-8" in text mode. Without it, open() falls back to the locale’s encoding, so the same file decodes one way on your Linux CI runner and another way on a teammate’s Windows laptop. Run your test suite with -X warn_default_encoding to catch the calls that forgot. AI coding assistants are especially prone to emitting a bare open(path) here.
  • Pick text or binary mode deliberately. A text file is decoded into str and gets universal newlines translation applied. A binary file hands you raw bytes. Images, archives, and anything you checksum belong in binary mode.
  • Stream large files instead of reading them whole. Iterating over a file object yields one line at a time and keeps memory flat no matter how big the file grows. Save .read() and .read_text() for files you know are small, like a config file.
  • Try the operation and handle the error instead of checking first. An os.path.exists() call before open() is already stale by the time the file opens, because another process can delete or create it in between. Catch FileNotFoundError or PermissionError instead, following the EAFP style that Python favors.
  • Replace files atomically when the old contents still matter. Opening a path in "w" mode truncates it right away, so a crash halfway through the write leaves you with neither the new data nor the old. Write to a temporary file in the same directory, then call os.replace(), which is atomic as long as both paths sit on one filesystem.
  • Use a format-aware module instead of writing your own parser. csv deals with quoting and embedded newlines, json deals with escaping, and shutil copies and archives whole trees. Open CSV files with newline="" so the module controls the line endings itself.
  • Never build a path straight out of untrusted input. A user-supplied name containing .. can climb out of the directory you meant to confine it to. Resolve the path first and confirm that it’s still inside the intended directory before you open it.

The atomic-replace rule is the easiest one to shrug off, because the damage only shows up on the run that gets interrupted. Move the power cut to any moment in the write and compare what each version leaves on disk:

Interactive diagram — enable JavaScript to view.

Writing in place passes through an empty file and then a half-written one, and a crash in that window destroys both versions. The temporary file plus os.replace() keeps the old contents readable right up to the instant they’re swapped for the new ones.

Say you need to count the error lines in a log file that may not exist yet. Compare two versions of the same function:

🔴 Avoid this:

Language: Python Filename: log_stats.py
import os

def count_errors(log_dir, name):
    path = log_dir + "/" + name
    if not os.path.exists(path):
        return 0
    log_file = open(path)
    lines = log_file.read().split("\n")
    return len([line for line in lines if "ERROR" in line])

This version works on the machine it was written on. The hardcoded / breaks on Windows, and the missing encoding argument makes the result depend on the local locale. On top of that, .read() pulls the entire log into memory, the os.path.exists() check can go stale before open() runs, and the file never gets closed.

✅ Favor this:

Language: Python Filename: log_stats.py
from pathlib import Path

def count_errors(log_dir, name):
    path = Path(log_dir) / name
    try:
        with path.open(encoding="utf-8") as log_file:
            return sum(1 for line in log_file if "ERROR" in line)
    except FileNotFoundError:
        return 0

This version builds the path with pathlib, spells out its encoding, and streams the file one line at a time, so memory use no longer tracks the size of the log. The with statement closes the file on the way out, and catching FileNotFoundError covers the missing file without a racy pre-check.

Reading and Writing Files in Python (Guide)

Tutorial

Reading and Writing Files in Python (Guide)

In this tutorial, you'll learn about reading and writing files in Python. You'll cover everything from what a file is made up of to which libraries can help you along that way. You'll also take a look at some basic scenarios of file usage as well as some advanced techniques.

intermediate python

For additional information on related topics, take a look at the following resources:

Have a question about this? Mentor AI can show you examples, compare related terms, and point you to tutorials.


By Martin Breuss • Updated Sept. 26, 2026