Capstone: analyze a bike-share dataset
Take a messy export from raw rows to a short, honest report.
- Clean a raw dataset with explicit, counted rules
- Summarize groups and patterns that answer a question
- Write findings that state scope and limits
The question: the city’s bike-share team asks, “How do members and casual riders use the bikes differently, and when is demand highest?” They’ve sent a raw export.
You’ll work like a data scientist would:
- Clean the export, counting how many rows each rule removes - those counts go in the report, because they affect the conclusions.
- Analyze the clean data: compare the rider groups, find the peak hour, and describe trip lengths.
- Report what you found, with sample sizes and caveats.
stations = [" Main St", "park ave", "HARBOR "]
print([station.strip().title() for station in stations])['Main St', 'Park Ave', 'Harbor']
Key takeaways
Clean with explicit rules and report how many rows each removed.
Answer the question with group comparisons and clear summaries.
State sample size, period and limitations alongside the findings.
Lesson quiz
5 questions · pass with 4 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Step 1: clean the export
Read the CSV (trip_id,start_hour,minutes,member,station). Apply, in order, counting removed rows:
- drop duplicate
trip_ids (keep the first) →duplicates: N; - treat
minutesof −1 as missing, then drop rows with missing minutes →missing minutes: N; - drop rows with minutes above 600 →
impossible minutes: N.
Then print kept K of T trips, clean station (strip spaces, title case) and print stations: NAME COUNT, ... alphabetically.
- Raw export
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Step 2: write the numbers for the report
Read a clean CSV (member,start_hour,minutes) and print:
1casual: n=N median=M min (one line per rider type, alphabetically; median to 1 decimal)
2member: n=N median=M min
3busiest hour: H (N trips) (the earliest hour if tied)
4short trips (<15 min): P% (whole percent)- Ten trips
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…