Python File Handling — Read, Write, and Manage Files Like a Pro

Sanjeev SharmaSanjeev Sharma
5 min read

Advertisement

Introduction

Why This Matters

Nearly every real-world Python program interacts with files — reading configuration, processing CSV data exports, writing logs, parsing JSON responses, or handling binary uploads. Python's file handling API is powerful, and knowing it well directly impacts your productivity in automation scripts, data pipelines, and backend services.

Python's pathlib module (introduced in Python 3.4) has modernized file operations, replacing string path manipulation with an object-oriented API. Combined with context managers (with statements), Python file handling is both safe and readable. Understanding these patterns is expected in data engineering, backend development, and DevOps scripting roles.

This guide covers text files, binary files, CSV, JSON, and the modern pathlib API.

Opening and Reading Files

# Basic read with context manager (auto-closes file)
with open("data.txt", "r", encoding="utf-8") as f:
    content = f.read()        # entire file as string
    print(content)
 
# Read line by line (memory efficient for large files)
with open("data.txt", "r", encoding="utf-8") as f:
    for line in f:
        print(line.strip())
 
# Read all lines into a list
with open("data.txt", "r") as f:
    lines = f.readlines()     # ['line1\n', 'line2\n', ...]
    lines = f.read().splitlines()  # ['line1', 'line2', ...] (no newlines)

Writing Files

# Write mode ('w') creates or overwrites
with open("output.txt", "w", encoding="utf-8") as f:
    f.write("Hello, World!\n")
    f.write("Second line\n")
 
# Append mode ('a') adds to existing file
with open("log.txt", "a", encoding="utf-8") as f:
    f.write("New log entry\n")
 
# Write multiple lines at once
lines = ["Line 1\n", "Line 2\n", "Line 3\n"]
with open("output.txt", "w") as f:
    f.writelines(lines)

File Modes Reference

ModeDescription
rRead (default); error if file not found
wWrite; creates or truncates file
aAppend; creates if not found
xExclusive create; error if file exists
rb, wbRead/write binary
r+Read and write

Binary Files

# Read an image file
with open("photo.jpg", "rb") as f:
    data = f.read()
    print(f"File size: {len(data)} bytes")
 
# Copy a binary file
with open("source.jpg", "rb") as src, open("copy.jpg", "wb") as dst:
    dst.write(src.read())
 
# Read in chunks (for large files)
def copy_large(src_path: str, dst_path: str, chunk_size: int = 8192):
    with open(src_path, "rb") as src, open(dst_path, "wb") as dst:
        while chunk := src.read(chunk_size):
            dst.write(chunk)

Working with CSV

import csv
 
# Write CSV
rows = [
    ["Name", "Age", "City"],
    ["Alice", 30, "New York"],
    ["Bob", 25, "London"],
]
with open("people.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.writer(f)
    writer.writerows(rows)
 
# Read CSV as dicts
with open("people.csv", "r", encoding="utf-8") as f:
    reader = csv.DictReader(f)
    for row in reader:
        print(row["Name"], row["Age"])  # Alice 30

Working with JSON

import json
 
data = {"name": "Alice", "scores": [95, 87, 92], "active": True}
 
# Write JSON
with open("data.json", "w", encoding="utf-8") as f:
    json.dump(data, f, indent=2)
 
# Read JSON
with open("data.json", "r", encoding="utf-8") as f:
    loaded = json.load(f)
    print(loaded["name"])  # Alice

pathlib — The Modern Way

from pathlib import Path
 
# Create path objects
base = Path("/tmp/myapp")
log_file = base / "logs" / "app.log"
 
# Check and create
base.mkdir(parents=True, exist_ok=True)
log_file.touch()
 
# Read and write
log_file.write_text("Starting application\n", encoding="utf-8")
content = log_file.read_text(encoding="utf-8")
 
# Glob and iterate
for py_file in Path(".").glob("**/*.py"):
    print(py_file.name, py_file.stat().st_size)
 
# File metadata
p = Path("data.txt")
if p.exists():
    print(p.suffix)       # '.txt'
    print(p.stem)         # 'data'
    print(p.parent)       # parent directory
    print(p.stat().st_size)  # size in bytes

Error Handling

from pathlib import Path
 
def safe_read(path: str) -> str:
    try:
        return Path(path).read_text(encoding="utf-8")
    except FileNotFoundError:
        print(f"File not found: {path}")
        return ""
    except PermissionError:
        print(f"No permission to read: {path}")
        return ""
    except UnicodeDecodeError:
        print(f"Cannot decode file: {path}")
        return ""

Common Mistakes

  • Forgetting to specify encoding="utf-8" — causes platform-specific encoding issues
  • Opening files without with — risk of unclosed file handles on exceptions
  • Using string concatenation for paths — path + "/file.txt" breaks on Windows
  • Reading large files with f.read() all at once — exhausts memory
  • Forgetting newline="" when writing CSV on Windows — causes double newlines

Best Practices

  • Always use with statements to ensure files are closed properly
  • Use pathlib.Path instead of os.path for modern, readable path handling
  • Specify encoding="utf-8" explicitly on every open() call
  • Read large files line by line or in chunks, not all at once
  • Use csv.DictReader and json.load() for structured data instead of parsing manually

Key Takeaways

  • The with open() context manager guarantees the file is closed even if an exception occurs
  • File modes: r (read), w (write/overwrite), a (append), rb/wb (binary)
  • pathlib.Path is the modern API for path operations — prefer it over os.path
  • csv.DictReader reads CSV rows as dictionaries keyed by header names
  • json.load(f) and json.dump(data, f, indent=2) handle JSON files cleanly
  • Always specify encoding="utf-8" to avoid platform-specific encoding bugs
  • Read large files in chunks or line by line to avoid memory exhaustion

Advertisement

Sanjeev Sharma

Written by

Sanjeev Sharma

Full Stack Engineer · E-mopro

Related reading