Curious about python for data analysis beginners and where to start? Start with pandas. Pandas is the library behind nearly every data job in Python — from startup dashboards to Netflix recommendations — and it is far more approachable than its reputation suggests. This pandas tutorial for beginners will take you from zero to analyzing a real dataset: creating DataFrames, reading CSV files, filtering rows, grouping data, and handling messy values, all with clear examples you can run yourself.
Why Python for Data Analysis Beginners Choose Pandas
Data analysis means turning raw numbers into answers: which product sells best, when do users sign up, what is the average score? Spreadsheets work for small data, but they break down at scale — and they cannot be automated. That is why python for data analysis beginners starts with pandas:
- Built for tables: pandas thinks in rows and columns, just like a spreadsheet, but handles millions of rows effortlessly.
- One-liners for big jobs: filtering, grouping, and summarizing thousands of records takes a single readable line.
- Plays well with others: pandas connects to NumPy, Matplotlib, scikit-learn, and SQL databases — the full data toolkit.
- Career fuel: “pandas” appears in most entry-level data analyst job posts. Learning it early pays off fast.
If you are new to Python itself, work through our Python automation scripts for beginners first to get comfortable with the language, or follow our complete coding roadmap from the very beginning.
Setting Up: Installing Pandas
You need Python 3 and the pandas library. Install it with pip:
pip install pandas
Verify the installation:
python -c "import pandas as pd; print(pd.__version__)"
The convention import pandas as pd is universal — every tutorial, book, and Stack Overflow answer uses pd, so adopt it now. The official pandas documentation is excellent when you want to go deeper on any function covered here.
Your First DataFrame: Pandas’ Core Object
In this pandas tutorial for beginners, everything revolves around the DataFrame — a table with labeled rows and columns. Create one from a plain Python dictionary:
import pandas as pd
data = {
"name": ["Ava", "Liam", "Mia", "Noah"],
"age": [28, 34, 29, 41],
"city": ["Austin", "Denver", "Austin", "Miami"],
"score": [88, 92, 79, 95],
}
df = pd.DataFrame(data)
print(df)
Output:
name age city score
0 Ava 28 Austin 88
1 Liam 34 Denver 92
2 Mia 29 Austin 79
3 Noah 41 Miami 95
Key ideas: each column has a name and a data type, and the numbers on the left are the index — row labels pandas manages for you. A single column is called a Series: df["age"] gives you just the ages.
Reading Real Data: CSV Files
Real-world data lives in CSV files. The read_csv function loads one into a DataFrame in a single line — the starting point of almost every python data analysis tutorial:
import pandas as pd
from io import StringIO
csv_text = """product,price,units_sold,region
laptop,999,120,North
mouse,25,850,North
keyboard,75,430,South
laptop,999,95,South
monitor,299,210,North
"""
df = pd.read_csv(StringIO(csv_text))
print(df)
(Here StringIO lets us embed the CSV directly so the example runs anywhere; normally you would write pd.read_csv("sales.csv") with a real file.) Pandas automatically detects the header row, infers that price is numeric, and builds the DataFrame — try doing that with raw Python and you will appreciate the one-liner.
Exploring Data: head, info, and describe
Before analyzing anything, get to know your data. These three commands are the first thing professionals run on any new dataset — memorize them:
print(df.head()) # first 5 rows
print(df.info()) # column names, types, non-null counts
print(df.describe()) # stats for numeric columns: mean, min, max...
describe() is pure gold for python for data analysis beginners: in one glance you see averages, spread, and extremes, which often reveal errors (a negative price?) before they corrupt your conclusions.
Selecting and Filtering Rows
Filtering is where pandas shines. Square brackets with a condition return only matching rows — no loops needed:
# all products that cost more than $100
expensive = df[df["price"] > 100]
print(expensive)
# products from the North region with more than 200 units sold
popular_north = df[(df["region"] == "North") & (df["units_sold"] > 200)]
print(popular_north)
# just the product and price columns, sorted by price
print(df[["product", "price"]].sort_values("price", ascending=False))
How it works: df["price"] > 100 produces a column of True/False values, and pandas keeps only the True rows — this is called boolean indexing. Combine conditions with & (and) or | (or), wrapping each in parentheses.
Grouping and Aggregating: Answering Real Questions
“What is the total revenue per region?” Doing that by hand means loops and running totals; with pandas it is one expressive line. This is the moment most learners truly learn pandas:
# add a revenue column first
df["revenue"] = df["price"] * df["units_sold"]
# total revenue and units per region
summary = df.groupby("region").agg(
total_revenue=("revenue", "sum"),
total_units=("units_sold", "sum"),
avg_price=("price", "mean"),
)
print(summary)
Output:
total_revenue total_units avg_price
region
North 265450 1180 441.0
South 127130 525 537.0
How it works: groupby("region") splits rows into North and South groups, then agg computes a summary per group. This single pattern — split, apply, combine — answers most business questions you will ever face, from “sales by month” to “average score by class”.
Handling Missing Data Like a Pro
Real datasets have holes — empty cells, “N/A” strings, corrupted rows. Pandas represents missing values as NaN (Not a Number) and gives you clean tools to deal with them:
import numpy as np
df2 = pd.DataFrame({
"name": ["Ava", "Liam", "Mia"],
"score": [88, np.nan, 79],
})
print(df2.isna().sum()) # count missing values per column
print(df2.dropna()) # remove rows with any missing value
filled = df2.fillna({"score": df2["score"].mean()})
print(filled) # fill gaps with the column average
Rule of thumb: dropna() is fine when few rows are affected; fillna() preserves data when dropping would lose too much. Either way, always check for missing values with isna().sum() before trusting any analysis — silent NaNs are the most common beginner trap in any python data analysis tutorial.
Mini Project: Analyze a Sales Dataset End to End
Let us put it all together. This complete python data analysis tutorial script loads sales data, cleans it, and answers three business questions — the exact workflow of a junior data analyst:
import pandas as pd
from io import StringIO
csv_text = """date,product,price,units,region
2026-10-01,laptop,999,12,North
2026-10-02,mouse,25,80,North
2026-10-03,keyboard,75,40,South
2026-10-04,laptop,999,9,South
2026-10-05,monitor,299,21,North
2026-10-06,mouse,25,,South
"""
df = pd.read_csv(StringIO(csv_text))
# 1. Clean: fill the missing units value with 0
df["units"] = df["units"].fillna(0)
# 2. Which product earned the most?
df["revenue"] = df["price"] * df["units"]
print("Revenue by product:")
print(df.groupby("product")["revenue"].sum()
.sort_values(ascending=False))
# 3. Which region sold more units?
print("\nUnits by region:")
print(df.groupby("region")["units"].sum())
# 4. Best sales day?
daily = df.groupby("date")["revenue"].sum()
print("\nBest day:", daily.idxmax(), f"(${daily.max():,.0f})")
Eight lines of logic replace what would be fifty lines of plain Python — that leverage is why companies pay analysts who learn pandas well.
Common Beginner Mistakes to Avoid
- Chained assignment warnings: writing
df[df["x"] > 1]["y"] = 5may silently fail. Usedf.loc[df["x"] > 1, "y"] = 5instead —.locis the safe way to set values on filtered rows. - Forgetting CSV headers: if your file has no header row, pass
header=Nonetoread_csv, or your first data row becomes column names. - Ignoring data types: numbers stored as text will not sum correctly. Check
df.dtypesand convert withpd.to_numeric()when needed. - Analyzing before cleaning: always run
isna().sum()anddescribe()first. Garbage in, garbage out.
What’s Next
You have just completed a genuine pandas tutorial for beginners: DataFrames, CSV loading, exploration, filtering, grouping, and missing-data handling — the core toolkit of python for data analysis beginners. To keep climbing, pair pandas with SQL (our SQL tutorial for beginners covers the queries analysts use daily), automate your data pipelines with our Python automation scripts, and build portfolio pieces from our Python projects with source code. For the full learning path, see how to learn coding from scratch in 2026.