Python File Handling — Read, Write, and Manage Files Like a Pro
Advertisement
Introduction
Why This Matters
Nearly every real-world Python program interacts with files — reading configuration, processing CSV data exports, writing logs, parsing JSON responses, or handling binary uploads. Python's file handling API is powerful, and knowing it well directly impacts your productivity in automation scripts, data pipelines, and backend services.
Python's pathlib module (introduced in Python 3.4) has modernized file operations, replacing string path manipulation with an object-oriented API. Combined with context managers (with statements), Python file handling is both safe and readable. Understanding these patterns is expected in data engineering, backend development, and DevOps scripting roles.
This guide covers text files, binary files, CSV, JSON, and the modern pathlib API.
Opening and Reading Files
# Basic read with context manager (auto-closes file)
with open("data.txt", "r", encoding="utf-8") as f:
content = f.read() # entire file as string
print(content)
# Read line by line (memory efficient for large files)
with open("data.txt", "r", encoding="utf-8") as f:
for line in f:
print(line.strip())
# Read all lines into a list
with open("data.txt", "r") as f:
lines = f.readlines() # ['line1\n', 'line2\n', ...]
lines = f.read().splitlines() # ['line1', 'line2', ...] (no newlines)Writing Files
# Write mode ('w') creates or overwrites
with open("output.txt", "w", encoding="utf-8") as f:
f.write("Hello, World!\n")
f.write("Second line\n")
# Append mode ('a') adds to existing file
with open("log.txt", "a", encoding="utf-8") as f:
f.write("New log entry\n")
# Write multiple lines at once
lines = ["Line 1\n", "Line 2\n", "Line 3\n"]
with open("output.txt", "w") as f:
f.writelines(lines)File Modes Reference
| Mode | Description |
|---|---|
r | Read (default); error if file not found |
w | Write; creates or truncates file |
a | Append; creates if not found |
x | Exclusive create; error if file exists |
rb, wb | Read/write binary |
r+ | Read and write |
Binary Files
# Read an image file
with open("photo.jpg", "rb") as f:
data = f.read()
print(f"File size: {len(data)} bytes")
# Copy a binary file
with open("source.jpg", "rb") as src, open("copy.jpg", "wb") as dst:
dst.write(src.read())
# Read in chunks (for large files)
def copy_large(src_path: str, dst_path: str, chunk_size: int = 8192):
with open(src_path, "rb") as src, open(dst_path, "wb") as dst:
while chunk := src.read(chunk_size):
dst.write(chunk)Working with CSV
import csv
# Write CSV
rows = [
["Name", "Age", "City"],
["Alice", 30, "New York"],
["Bob", 25, "London"],
]
with open("people.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.writer(f)
writer.writerows(rows)
# Read CSV as dicts
with open("people.csv", "r", encoding="utf-8") as f:
reader = csv.DictReader(f)
for row in reader:
print(row["Name"], row["Age"]) # Alice 30Working with JSON
import json
data = {"name": "Alice", "scores": [95, 87, 92], "active": True}
# Write JSON
with open("data.json", "w", encoding="utf-8") as f:
json.dump(data, f, indent=2)
# Read JSON
with open("data.json", "r", encoding="utf-8") as f:
loaded = json.load(f)
print(loaded["name"]) # Alicepathlib — The Modern Way
from pathlib import Path
# Create path objects
base = Path("/tmp/myapp")
log_file = base / "logs" / "app.log"
# Check and create
base.mkdir(parents=True, exist_ok=True)
log_file.touch()
# Read and write
log_file.write_text("Starting application\n", encoding="utf-8")
content = log_file.read_text(encoding="utf-8")
# Glob and iterate
for py_file in Path(".").glob("**/*.py"):
print(py_file.name, py_file.stat().st_size)
# File metadata
p = Path("data.txt")
if p.exists():
print(p.suffix) # '.txt'
print(p.stem) # 'data'
print(p.parent) # parent directory
print(p.stat().st_size) # size in bytesError Handling
from pathlib import Path
def safe_read(path: str) -> str:
try:
return Path(path).read_text(encoding="utf-8")
except FileNotFoundError:
print(f"File not found: {path}")
return ""
except PermissionError:
print(f"No permission to read: {path}")
return ""
except UnicodeDecodeError:
print(f"Cannot decode file: {path}")
return ""Common Mistakes
- Forgetting to specify
encoding="utf-8"— causes platform-specific encoding issues - Opening files without
with— risk of unclosed file handles on exceptions - Using string concatenation for paths —
path + "/file.txt"breaks on Windows - Reading large files with
f.read()all at once — exhausts memory - Forgetting
newline=""when writing CSV on Windows — causes double newlines
Best Practices
- Always use
withstatements to ensure files are closed properly - Use
pathlib.Pathinstead ofos.pathfor modern, readable path handling - Specify
encoding="utf-8"explicitly on everyopen()call - Read large files line by line or in chunks, not all at once
- Use
csv.DictReaderandjson.load()for structured data instead of parsing manually
Key Takeaways
- The
with open()context manager guarantees the file is closed even if an exception occurs - File modes:
r(read),w(write/overwrite),a(append),rb/wb(binary) pathlib.Pathis the modern API for path operations — prefer it overos.pathcsv.DictReaderreads CSV rows as dictionaries keyed by header namesjson.load(f)andjson.dump(data, f, indent=2)handle JSON files cleanly- Always specify
encoding="utf-8"to avoid platform-specific encoding bugs - Read large files in chunks or line by line to avoid memory exhaustion
Advertisement