Cheat sheetPython for Data ScienceNumPy, pandas, cleaning, reshaping, time series, plotting, scikit-learn pipelines, and performance habits. 2]` selects elements without a loop."}],"id":"bl_04_004"}]},{"id":"card_04_002","type":"card","x":0,"y":1008,"w":560,"h":462,"z":13,"title":"NumPy Math","sectionNumber":"3","accent":"accent","bg":"surface","showHeader":true,"blocks":[{"type":"code","language":"python","code":"a.sum(axis=0) # column sums\na.mean(axis=1) # row means\na @ b # matrix product\nnp.where(a > 2, 1, 0)\nnp.clip(a, 0, 5)\nnp.percentile(a, [25, 50, 75])","id":"bl_04_005"},{"type":"keyvalue","items":[{"term":"NaN-aware","desc":"`np.nanmean`, `np.nansum` skip missing values."},{"term":"Random","desc":"`rng = np.random.default_rng(42)` for reproducible draws."}],"id":"bl_04_006"},{"type":"callout","variant":"tip","title":"Think in axes","text":"axis=0 collapses rows (one result per column); axis=1 collapses columns (one result per row).","id":"bl_04_007"}]},{"id":"card_04_003","type":"card","x":0,"y":1494,"w":560,"h":373,"z":14,"title":"pandas Basics","sectionNumber":"4","accent":"accent","bg":"surface","showHeader":true,"blocks":[{"type":"code","language":"python","code":"df = pd.read_csv(\"orders.csv\", parse_dates=[\"ts\"])\ndf.head(); df.info(); df.describe()\ndf[\"revenue\"].value_counts(dropna=False)\ndf.loc[df[\"country\"] == \"IN\", [\"user_id\", \"revenue\"]]\ndf.iloc[:5, :3]\ndf.assign(gross=lambda d: d[\"revenue\"] * 1.18)","id":"bl_04_008"},{"type":"keyvalue","items":[{"term":"loc vs iloc","desc":"`loc` uses labels and boolean masks; `iloc` uses integer positions."},{"term":"Chaining","desc":"Prefer `.assign()` and `.pipe()` to avoid SettingWithCopyWarning."}],"id":"bl_04_009"}]},{"id":"card_04_004","type":"card","x":592,"y":137,"w":560,"h":461,"z":15,"title":"Missing Data","sectionNumber":"5","accent":"accent","bg":"surface","showHeader":true,"blocks":[{"type":"code","language":"python","code":"df.isna().mean().sort_values(ascending=False)\ndf.dropna(subset=[\"user_id\"])\ndf[\"age\"] = df[\"age\"].fillna(df[\"age\"].median())\ndf[\"city\"] = df[\"city\"].fillna(\"unknown\")\ndf[\"age_missing\"] = df[\"age\"].isna()","id":"bl_04_00a"},{"type":"keyvalue","items":[{"term":"Sentinels","desc":"Convert placeholders like -1, \"N/A\", or \"\" to `NaN` first."},{"term":"Indicator column","desc":"Keep \"was missing\" when absence itself carries signal."}],"id":"bl_04_00b"},{"type":"callout","variant":"warn","title":"Impute after splitting","text":"Compute medians and modes on the training data only, then apply them to validation and test.","id":"bl_04_00c"}]},{"id":"card_04_005","type":"card","x":592,"y":622,"w":560,"h":387,"z":16,"title":"Filter, Sort, Deduplicate","sectionNumber":"6","accent":"accent","bg":"surface","showHeader":true,"blocks":[{"type":"code","language":"python","code":"df.query(\"revenue > 100 and status == 'paid'\")\ndf[df[\"country\"].isin([\"IN\", \"US\"])]\ndf.sort_values([\"user_id\", \"ts\"], ascending=[True, False])\ndf.drop_duplicates(subset=[\"order_id\"], keep=\"last\")\ndf.nlargest(10, \"revenue\")","id":"bl_04_00d"},{"type":"keyvalue","items":[{"term":"query","desc":"Readable filters; reference variables with `@name`."},{"term":"keep","desc":"`\"first\"`, `\"last\"`, or `False` to drop every duplicate."},{"term":"Stable sort","desc":"Pass `kind=\"stable\"` when ties must keep their order."}],"id":"bl_04_00e"}]},{"id":"card_04_006","type":"card","x":592,"y":1033,"w":560,"h":389,"z":17,"title":"GroupBy and Aggregation","sectionNumber":"7","accent":"accent","bg":"surface","showHeader":true,"blocks":[{"type":"code","language":"python","code":"(df.groupby(\"country\", as_index=False)\n .agg(users=(\"user_id\", \"nunique\"),\n revenue=(\"revenue\", \"sum\"),\n aov=(\"revenue\", \"mean\")))\ndf[\"share\"] = df[\"revenue\"] / df.groupby(\n \"country\")[\"revenue\"].transform(\"sum\")","id":"bl_04_00f"},{"type":"keyvalue","items":[{"term":"agg","desc":"One row per group."},{"term":"transform","desc":"Same shape as the input; good for group-wise features."},{"term":"filter","desc":"Keep whole groups that satisfy a condition."}],"id":"bl_04_00g"}]},{"id":"card_04_007","type":"card","x":592,"y":1446,"w":560,"h":514,"z":18,"title":"Merge and Join","sectionNumber":"8","accent":"accent","bg":"surface","showHeader":true,"blocks":[{"type":"code","language":"python","code":"orders.merge(users, on=\"user_id\", how=\"left\",\n validate=\"many_to_one\", indicator=True)\npd.concat([df_2025, df_2026], ignore_index=True)","id":"bl_04_00h"},{"type":"table","headers":["how","Keeps"],"rows":[["inner","Keys present in both"],["left","All left rows"],["outer","All keys from both"],["cross","Every combination"]],"id":"bl_04_00i"},{"type":"callout","variant":"warn","title":"Validate join cardinality","text":"`validate=\"many_to_one\"` raises if the right key is duplicated, catching silent row explosions.","id":"bl_04_00j"}]},{"id":"card_04_008","type":"card","x":1184,"y":137,"w":560,"h":429,"z":19,"title":"Reshaping","sectionNumber":"9","accent":"accent","bg":"surface","showHeader":true,"blocks":[{"type":"code","language":"python","code":"wide = df.pivot_table(index=\"user_id\",\n columns=\"month\",\n values=\"revenue\",\n aggfunc=\"sum\", fill_value=0)\nlong = wide.reset_index().melt(id_vars=\"user_id\",\n var_name=\"month\")\ndf.explode(\"tags\")","id":"bl_04_00k"},{"type":"keyvalue","items":[{"term":"pivot_table","desc":"Long → wide with aggregation; handles duplicate keys."},{"term":"melt","desc":"Wide → long; the tidy shape most plotting and modeling tools expect."},{"term":"explode","desc":"One row per list element."}],"id":"bl_04_00l"}]},{"id":"card_04_009","type":"card","x":1184,"y":590,"w":560,"h":352,"z":20,"title":"Dates and Time Series","sectionNumber":"10","accent":"accent","bg":"surface","showHeader":true,"blocks":[{"type":"code","language":"python","code":"df[\"ts\"] = pd.to_datetime(df[\"ts\"], utc=True)\ndf[\"dow\"] = df[\"ts\"].dt.dayofweek\ndaily = df.set_index(\"ts\").resample(\"D\")[\"revenue\"].sum()\ndaily.rolling(7).mean()\ndf[\"prev\"] = df.groupby(\"user_id\")[\"revenue\"].shift(1)","id":"bl_04_00m"},{"type":"keyvalue","items":[{"term":"Time zones","desc":"Store UTC; convert with `.dt.tz_convert()` for reporting."},{"term":"Gaps","desc":"`resample` creates missing periods; decide between fill with 0 and NaN."}],"id":"bl_04_00n"}]},{"id":"card_04_00a","type":"card","x":1184,"y":966,"w":560,"h":374,"z":21,"title":"Vectorize, Don't Loop","sectionNumber":"11","accent":"accent","bg":"surface","showHeader":true,"blocks":[{"type":"code","language":"python","code":"# slow: row-wise Python\ndf[\"band\"] = df[\"revenue\"].apply(\n lambda r: \"hi\" if r > 100 else \"lo\")\n# fast: vectorized\ndf[\"band\"] = np.where(df[\"revenue\"] > 100, \"hi\", \"lo\")\ndf[\"band\"] = pd.cut(df[\"revenue\"], [0, 100, np.inf],\n labels=[\"lo\", \"hi\"])","id":"bl_04_00o"},{"type":"keyvalue","items":[{"term":"map","desc":"Element-wise lookup from a dict or Series."},{"term":"Categoricals","desc":"`astype(\"category\")` saves memory for repeated strings."}],"id":"bl_04_00p"}]},{"id":"card_04_00b","type":"card","x":1184,"y":1364,"w":560,"h":421,"z":22,"title":"Files and Formats","sectionNumber":"12","accent":"accent","bg":"surface","showHeader":true,"blocks":[{"type":"keyvalue","items":[{"term":"CSV","desc":"Human-readable; slow and loses dtypes. Pass `dtype` and `parse_dates`."},{"term":"Parquet","desc":"Columnar, typed, compressed; the default for analytics."},{"term":"Feather","desc":"Fast local interchange between Python and R."},{"term":"Chunks","desc":"`pd.read_csv(path, chunksize=100_000)` streams large files."}],"id":"bl_04_00q"},{"type":"code","language":"python","code":"df.to_parquet(\"orders.parquet\", index=False)\ndf = pd.read_parquet(\"orders.parquet\",\n columns=[\"user_id\", \"revenue\"])","id":"bl_04_00r"}]},{"id":"card_04_00c","type":"card","x":1184,"y":1809,"w":560,"h":474,"z":23,"title":"Plotting","sectionNumber":"13","accent":"accent","bg":"surface","showHeader":true,"blocks":[{"type":"code","language":"python","code":"fig, ax = plt.subplots(figsize=(8, 4))\ndaily.plot(ax=ax, label=\"daily\")\ndaily.rolling(7).mean().plot(ax=ax, label=\"7-day\")\nax.set(title=\"Revenue\", xlabel=\"\", ylabel=\"USD\")\nax.legend(); fig.tight_layout()\nfig.savefig(\"revenue.png\", dpi=200)","id":"bl_04_00s"},{"type":"table","headers":["Question","Chart"],"rows":[["Trend over time","Line"],["Compare categories","Bar"],["Distribution","Histogram or box plot"],["Relationship","Scatter"]],"id":"bl_04_00t"}]},{"id":"card_04_00d","type":"card","x":1776,"y":137,"w":560,"h":408,"z":24,"title":"Data Quality Checks","sectionNumber":"14","accent":"accent","bg":"surface","showHeader":true,"blocks":[{"type":"code","language":"python","code":"assert df[\"order_id\"].is_unique\nassert df[\"revenue\"].ge(0).all()\nassert df[\"ts\"].between(\"2020-01-01\", \"2030-01-01\").all()\ndf.duplicated().sum()\ndf.select_dtypes(\"number\").describe().T","id":"bl_04_00u"},{"type":"keyvalue","items":[{"term":"Keys","desc":"Unique ids where expected; no orphan foreign keys after joins."},{"term":"Ranges","desc":"Non-negative amounts, plausible dates, valid categories."},{"term":"Freshness","desc":"The latest timestamp is as recent as the pipeline promises."}],"id":"bl_04_00v"}]},{"id":"card_04_00e","type":"card","x":1776,"y":569,"w":560,"h":586,"z":25,"title":"scikit-learn Pipeline","sectionNumber":"15","accent":"accent","bg":"surface","showHeader":true,"blocks":[{"type":"code","language":"python","code":"from sklearn.pipeline import Pipeline\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.preprocessing import OneHotEncoder\nfrom sklearn.linear_model import LogisticRegression\n\npre = ColumnTransformer([\n (\"num\", StandardScaler(), num_cols),\n (\"cat\", OneHotEncoder(handle_unknown=\"ignore\"),\n cat_cols),\n])\nmodel = Pipeline([(\"pre\", pre),\n (\"clf\", LogisticRegression())])\nmodel.fit(X_train, y_train)","id":"bl_04_00w"},{"type":"callout","variant":"tip","title":"Pipelines prevent leakage","text":"Fitting the whole pipeline inside cross-validation keeps scalers and encoders train-only.","id":"bl_04_00x"},{"type":"keyvalue","items":[{"term":"handle_unknown=\"ignore\"","desc":"New categories at inference become all-zero instead of an error."}],"id":"bl_04_00y"}]},{"id":"card_04_00f","type":"card","x":1776,"y":1179,"w":560,"h":395,"z":26,"title":"Evaluation Snippets","sectionNumber":"16","accent":"accent","bg":"surface","showHeader":true,"blocks":[{"type":"code","language":"python","code":"from sklearn.model_selection import cross_val_score\nfrom sklearn.metrics import classification_report\n\nscores = cross_val_score(model, X, y, cv=5,\n scoring=\"roc_auc\")\nprint(scores.mean(), scores.std())\ny_pred = model.predict(X_test)\nprint(classification_report(y_test, y_pred))","id":"bl_04_00z"},{"type":"keyvalue","items":[{"term":"Report the spread","desc":"Mean ± std across folds, not one lucky split."},{"term":"Probabilities","desc":"Use `predict_proba` for ROC-AUC, PR-AUC, and threshold tuning."}],"id":"bl_04_010"}]},{"id":"card_04_00g","type":"card","x":1776,"y":1598,"w":560,"h":348,"z":27,"title":"Performance Tips","sectionNumber":"17","accent":"accent","bg":"surface","showHeader":true,"blocks":[{"type":"keyvalue","items":[{"term":"Dtypes","desc":"Downcast numbers and use categories for repeated strings."},{"term":"Read less","desc":"Load only the needed columns and rows."},{"term":"Avoid apply","desc":"Vectorized NumPy/pandas operations are 10–100× faster."},{"term":"Profile","desc":"`%timeit` and `df.memory_usage(deep=True)` find the real hot spots."},{"term":"Scale out","desc":"Polars or DuckDB handle larger-than-pandas workloads on one machine."}],"id":"bl_04_011"}]},{"id":"card_04_00h","type":"card","x":1776,"y":1970,"w":560,"h":235,"z":28,"title":"Debug Checklist","sectionNumber":"18","accent":"accent","bg":"surface","showHeader":true,"blocks":[{"type":"list","ordered":true,"items":["Check `df.shape` after every merge and filter.","Look at `dtypes`; numbers stored as strings break math silently.","Count nulls before and after each transform.","Verify keys are unique where you expect them to be.","Set seeds (`random_state`) for reproducible splits and models."],"id":"bl_04_012"}]}],"background":"dots","canvasColor":"canvas","themeId":"midnight","layoutGap":24,"updatedAt":"2026-10-02T09:26:37.557Z"}]]>Python for Data Science NumPy, pandas, cleaning, reshaping, time series, plotting, scikit-learn pipelines, and performance habits. 01Workflow and Imports1Load2Inspect3Clean4Transform5Aggregate6Model or plotKeep each step a small, testable functionPythonimport numpy as np import pandas as pd import matplotlib.pyplot as plt from sklearn.model_selection import train_test_splitNotebooksRestart and run all before trusting a result; hidden state causes bugs.EnvironmentsPin versions in requirements.txt or pyproject.toml.02NumPy ArraysPythona = np.array([[1, 2, 3], [4, 5, 6]]) a.shape, a.ndim, a.dtype # (2, 3), 2, int64 np.zeros((3, 4)); np.ones(5); np.arange(0, 10, 2) np.linspace(0, 1, 5) # 5 evenly spaced a.reshape(3, 2); a.T; a.ravel()Views vs copiesSlices are views; use .copy() before mutating a subset.BroadcastingShapes align from the right; size-1 axes stretch to match.Boolean masksa[a > 2] selects elements without a loop.03NumPy MathPythona.sum(axis=0) # column sums a.mean(axis=1) # row means a @ b # matrix product np.where(a > 2, 1, 0) np.clip(a, 0, 5) np.percentile(a, [25, 50, 75])NaN-awarenp.nanmean, np.nansum skip missing values.Randomrng = np.random.default_rng(42) for reproducible draws.Think in axesaxis=0 collapses rows (one result per column); axis=1 collapses columns (one result per row). 04pandas BasicsPythondf = pd.read_csv("orders.csv", parse_dates=["ts"]) df.head(); df.info(); df.describe() df["revenue"].value_counts(dropna=False) df.loc[df["country"] == "IN", ["user_id", "revenue"]] df.iloc[:5, :3] df.assign(gross=lambda d: d["revenue"] * 1.18)loc vs ilocloc uses labels and boolean masks; iloc uses integer positions.ChainingPrefer .assign() and .pipe() to avoid SettingWithCopyWarning.05Missing DataPythondf.isna().mean().sort_values(ascending=False) df.dropna(subset=["user_id"]) df["age"] = df["age"].fillna(df["age"].median()) df["city"] = df["city"].fillna("unknown") df["age_missing"] = df["age"].isna()SentinelsConvert placeholders like -1, "N/A", or "" to NaN first.Indicator columnKeep "was missing" when absence itself carries signal.Impute after splittingCompute medians and modes on the training data only, then apply them to validation and test. 06Filter, Sort, DeduplicatePythondf.query("revenue > 100 and status == 'paid'") df[df["country"].isin(["IN", "US"])] df.sort_values(["user_id", "ts"], ascending=[True, False]) df.drop_duplicates(subset=["order_id"], keep="last") df.nlargest(10, "revenue")queryReadable filters; reference variables with @name.keep"first", "last", or False to drop every duplicate.Stable sortPass kind="stable" when ties must keep their order.07GroupBy and AggregationPython(df.groupby("country", as_index=False) .agg(users=("user_id", "nunique"), revenue=("revenue", "sum"), aov=("revenue", "mean"))) df["share"] = df["revenue"] / df.groupby( "country")["revenue"].transform("sum")aggOne row per group.transformSame shape as the input; good for group-wise features.filterKeep whole groups that satisfy a condition.08Merge and JoinPythonorders.merge(users, on="user_id", how="left", validate="many_to_one", indicator=True) pd.concat([df_2025, df_2026], ignore_index=True)howKeepsinnerKeys present in bothleftAll left rowsouterAll keys from bothcrossEvery combinationValidate join cardinalityvalidate="many_to_one" raises if the right key is duplicated, catching silent row explosions. 09ReshapingPythonwide = df.pivot_table(index="user_id", columns="month", values="revenue", aggfunc="sum", fill_value=0) long = wide.reset_index().melt(id_vars="user_id", var_name="month") df.explode("tags")pivot_tableLong → wide with aggregation; handles duplicate keys.meltWide → long; the tidy shape most plotting and modeling tools expect.explodeOne row per list element.10Dates and Time SeriesPythondf["ts"] = pd.to_datetime(df["ts"], utc=True) df["dow"] = df["ts"].dt.dayofweek daily = df.set_index("ts").resample("D")["revenue"].sum() daily.rolling(7).mean() df["prev"] = df.groupby("user_id")["revenue"].shift(1)Time zonesStore UTC; convert with .dt.tz_convert() for reporting.Gapsresample creates missing periods; decide between fill with 0 and NaN.11Vectorize, Don't LoopPython# slow: row-wise Python df["band"] = df["revenue"].apply( lambda r: "hi" if r > 100 else "lo") # fast: vectorized df["band"] = np.where(df["revenue"] > 100, "hi", "lo") df["band"] = pd.cut(df["revenue"], [0, 100, np.inf], labels=["lo", "hi"])mapElement-wise lookup from a dict or Series.Categoricalsastype("category") saves memory for repeated strings.12Files and FormatsCSVHuman-readable; slow and loses dtypes. Pass dtype and parse_dates.ParquetColumnar, typed, compressed; the default for analytics.FeatherFast local interchange between Python and R.Chunkspd.read_csv(path, chunksize=100_000) streams large files.Pythondf.to_parquet("orders.parquet", index=False) df = pd.read_parquet("orders.parquet", columns=["user_id", "revenue"])13PlottingPythonfig, ax = plt.subplots(figsize=(8, 4)) daily.plot(ax=ax, label="daily") daily.rolling(7).mean().plot(ax=ax, label="7-day") ax.set(title="Revenue", xlabel="", ylabel="USD") ax.legend(); fig.tight_layout() fig.savefig("revenue.png", dpi=200)QuestionChartTrend over timeLineCompare categoriesBarDistributionHistogram or box plotRelationshipScatter14Data Quality ChecksPythonassert df["order_id"].is_unique assert df["revenue"].ge(0).all() assert df["ts"].between("2020-01-01", "2030-01-01").all() df.duplicated().sum() df.select_dtypes("number").describe().TKeysUnique ids where expected; no orphan foreign keys after joins.RangesNon-negative amounts, plausible dates, valid categories.FreshnessThe latest timestamp is as recent as the pipeline promises.15scikit-learn PipelinePythonfrom sklearn.pipeline import Pipeline from sklearn.compose import ColumnTransformer from sklearn.preprocessing import StandardScaler from sklearn.preprocessing import OneHotEncoder from sklearn.linear_model import LogisticRegression pre = ColumnTransformer([ ("num", StandardScaler(), num_cols), ("cat", OneHotEncoder(handle_unknown="ignore"), cat_cols), ]) model = Pipeline([("pre", pre), ("clf", LogisticRegression())]) model.fit(X_train, y_train)Pipelines prevent leakageFitting the whole pipeline inside cross-validation keeps scalers and encoders train-only. handle_unknown="ignore"New categories at inference become all-zero instead of an error.16Evaluation SnippetsPythonfrom sklearn.model_selection import cross_val_score from sklearn.metrics import classification_report scores = cross_val_score(model, X, y, cv=5, scoring="roc_auc") print(scores.mean(), scores.std()) y_pred = model.predict(X_test) print(classification_report(y_test, y_pred))Report the spreadMean ± std across folds, not one lucky split.ProbabilitiesUse predict_proba for ROC-AUC, PR-AUC, and threshold tuning.17Performance TipsDtypesDowncast numbers and use categories for repeated strings.Read lessLoad only the needed columns and rows.Avoid applyVectorized NumPy/pandas operations are 10–100× faster.Profile%timeit and df.memory_usage(deep=True) find the real hot spots.Scale outPolars or DuckDB handle larger-than-pandas workloads on one machine.18Debug Checklist1.Check df.shape after every merge and filter.2.Look at dtypes; numbers stored as strings break math silently.3.Count nulls before and after each transform.4.Verify keys are unique where you expect them to be.5.Set seeds (random_state) for reproducible splits and models.