Python for Data Analysis: A Beginner’s Guide to Pandas in 2026

Python for data analysis beginners: learn pandas from scratch in 2026. DataFrames, CSV loading, filtering, groupby and a full mini project with code.

Curious about python for data analysis beginners and where to start? Start with pandas. Pandas is the library behind nearly every data job in Python — from startup dashboards to Netflix recommendations — and it is far more approachable than its reputation suggests. This pandas tutorial for beginners will take you from zero to analyzing a real dataset: creating DataFrames, reading CSV files, filtering rows, grouping data, and handling messy values, all with clear examples you can run yourself.

Why Python for Data Analysis Beginners Choose Pandas

Data analysis means turning raw numbers into answers: which product sells best, when do users sign up, what is the average score? Spreadsheets work for small data, but they break down at scale — and they cannot be automated. That is why python for data analysis beginners starts with pandas:

  • Built for tables: pandas thinks in rows and columns, just like a spreadsheet, but handles millions of rows effortlessly.
  • One-liners for big jobs: filtering, grouping, and summarizing thousands of records takes a single readable line.
  • Plays well with others: pandas connects to NumPy, Matplotlib, scikit-learn, and SQL databases — the full data toolkit.
  • Career fuel: “pandas” appears in most entry-level data analyst job posts. Learning it early pays off fast.

If you are new to Python itself, work through our Python automation scripts for beginners first to get comfortable with the language, or follow our complete coding roadmap from the very beginning.

Setting Up: Installing Pandas

You need Python 3 and the pandas library. Install it with pip:

pip install pandas

Verify the installation:

python -c "import pandas as pd; print(pd.__version__)"

The convention import pandas as pd is universal — every tutorial, book, and Stack Overflow answer uses pd, so adopt it now. The official pandas documentation is excellent when you want to go deeper on any function covered here.

Your First DataFrame: Pandas’ Core Object

In this pandas tutorial for beginners, everything revolves around the DataFrame — a table with labeled rows and columns. Create one from a plain Python dictionary:

import pandas as pd

data = {
    "name": ["Ava", "Liam", "Mia", "Noah"],
    "age": [28, 34, 29, 41],
    "city": ["Austin", "Denver", "Austin", "Miami"],
    "score": [88, 92, 79, 95],
}

df = pd.DataFrame(data)
print(df)

Output:

   name  age    city  score
0   Ava   28  Austin     88
1  Liam   34  Denver     92
2   Mia   29  Austin     79
3  Noah   41   Miami     95

Key ideas: each column has a name and a data type, and the numbers on the left are the index — row labels pandas manages for you. A single column is called a Series: df["age"] gives you just the ages.

Reading Real Data: CSV Files

Real-world data lives in CSV files. The read_csv function loads one into a DataFrame in a single line — the starting point of almost every python data analysis tutorial:

import pandas as pd
from io import StringIO

csv_text = """product,price,units_sold,region
laptop,999,120,North
mouse,25,850,North
keyboard,75,430,South
laptop,999,95,South
monitor,299,210,North
"""

df = pd.read_csv(StringIO(csv_text))
print(df)

(Here StringIO lets us embed the CSV directly so the example runs anywhere; normally you would write pd.read_csv("sales.csv") with a real file.) Pandas automatically detects the header row, infers that price is numeric, and builds the DataFrame — try doing that with raw Python and you will appreciate the one-liner.

Exploring Data: head, info, and describe

Before analyzing anything, get to know your data. These three commands are the first thing professionals run on any new dataset — memorize them:

print(df.head())      # first 5 rows
print(df.info())      # column names, types, non-null counts
print(df.describe())  # stats for numeric columns: mean, min, max...

describe() is pure gold for python for data analysis beginners: in one glance you see averages, spread, and extremes, which often reveal errors (a negative price?) before they corrupt your conclusions.

Selecting and Filtering Rows

Filtering is where pandas shines. Square brackets with a condition return only matching rows — no loops needed:

# all products that cost more than $100
expensive = df[df["price"] > 100]
print(expensive)

# products from the North region with more than 200 units sold
popular_north = df[(df["region"] == "North") & (df["units_sold"] > 200)]
print(popular_north)

# just the product and price columns, sorted by price
print(df[["product", "price"]].sort_values("price", ascending=False))

How it works: df["price"] > 100 produces a column of True/False values, and pandas keeps only the True rows — this is called boolean indexing. Combine conditions with & (and) or | (or), wrapping each in parentheses.

Grouping and Aggregating: Answering Real Questions

“What is the total revenue per region?” Doing that by hand means loops and running totals; with pandas it is one expressive line. This is the moment most learners truly learn pandas:

# add a revenue column first
df["revenue"] = df["price"] * df["units_sold"]

# total revenue and units per region
summary = df.groupby("region").agg(
    total_revenue=("revenue", "sum"),
    total_units=("units_sold", "sum"),
    avg_price=("price", "mean"),
)
print(summary)

Output:

        total_revenue  total_units  avg_price
region
North          265450         1180      441.0
South          127130          525      537.0

How it works: groupby("region") splits rows into North and South groups, then agg computes a summary per group. This single pattern — split, apply, combine — answers most business questions you will ever face, from “sales by month” to “average score by class”.

Handling Missing Data Like a Pro

Real datasets have holes — empty cells, “N/A” strings, corrupted rows. Pandas represents missing values as NaN (Not a Number) and gives you clean tools to deal with them:

import numpy as np

df2 = pd.DataFrame({
    "name": ["Ava", "Liam", "Mia"],
    "score": [88, np.nan, 79],
})

print(df2.isna().sum())   # count missing values per column
print(df2.dropna())       # remove rows with any missing value

filled = df2.fillna({"score": df2["score"].mean()})
print(filled)             # fill gaps with the column average

Rule of thumb: dropna() is fine when few rows are affected; fillna() preserves data when dropping would lose too much. Either way, always check for missing values with isna().sum() before trusting any analysis — silent NaNs are the most common beginner trap in any python data analysis tutorial.

Mini Project: Analyze a Sales Dataset End to End

Let us put it all together. This complete python data analysis tutorial script loads sales data, cleans it, and answers three business questions — the exact workflow of a junior data analyst:

import pandas as pd
from io import StringIO

csv_text = """date,product,price,units,region
2026-10-01,laptop,999,12,North
2026-10-02,mouse,25,80,North
2026-10-03,keyboard,75,40,South
2026-10-04,laptop,999,9,South
2026-10-05,monitor,299,21,North
2026-10-06,mouse,25,,South
"""

df = pd.read_csv(StringIO(csv_text))

# 1. Clean: fill the missing units value with 0
df["units"] = df["units"].fillna(0)

# 2. Which product earned the most?
df["revenue"] = df["price"] * df["units"]
print("Revenue by product:")
print(df.groupby("product")["revenue"].sum()
        .sort_values(ascending=False))

# 3. Which region sold more units?
print("\nUnits by region:")
print(df.groupby("region")["units"].sum())

# 4. Best sales day?
daily = df.groupby("date")["revenue"].sum()
print("\nBest day:", daily.idxmax(), f"(${daily.max():,.0f})")

Eight lines of logic replace what would be fifty lines of plain Python — that leverage is why companies pay analysts who learn pandas well.

Common Beginner Mistakes to Avoid

  • Chained assignment warnings: writing df[df["x"] > 1]["y"] = 5 may silently fail. Use df.loc[df["x"] > 1, "y"] = 5 instead — .loc is the safe way to set values on filtered rows.
  • Forgetting CSV headers: if your file has no header row, pass header=None to read_csv, or your first data row becomes column names.
  • Ignoring data types: numbers stored as text will not sum correctly. Check df.dtypes and convert with pd.to_numeric() when needed.
  • Analyzing before cleaning: always run isna().sum() and describe() first. Garbage in, garbage out.

What’s Next

You have just completed a genuine pandas tutorial for beginners: DataFrames, CSV loading, exploration, filtering, grouping, and missing-data handling — the core toolkit of python for data analysis beginners. To keep climbing, pair pandas with SQL (our SQL tutorial for beginners covers the queries analysts use daily), automate your data pipelines with our Python automation scripts, and build portfolio pieces from our Python projects with source code. For the full learning path, see how to learn coding from scratch in 2026.

Leave a Reply

Your email address will not be published. Required fields are marked *