• Home
  • About Us
  • Contact Us
  • Disclaimer
  • Privacy Policy
Thursday, August 13, 2026
newsaiworld
  • Home
  • Artificial Intelligence
  • ChatGPT
  • Data Science
  • Machine Learning
  • Crypto Coins
  • Contact Us
No Result
View All Result
  • Home
  • Artificial Intelligence
  • ChatGPT
  • Data Science
  • Machine Learning
  • Crypto Coins
  • Contact Us
No Result
View All Result
Morning News
No Result
View All Result
Home Data Science

Constructing an Finish-to-Finish Knowledge Science Portfolio Mission

Admin by Admin
August 13, 2026
in Data Science
0
Rosidi End to End Data Science Portfolio Project 1.png
0
SHARES
0
VIEWS
Share on FacebookShare on Twitter


Building End to End Data Science Portfolio Project
 

The initiatives that really get individuals employed do one thing totally different. They begin with a enterprise drawback and end with a suggestion, they usually present each stage in between. Exhibiting each stage, from uncooked knowledge to a deployed utility, is the factor a resume can’t show and a pocket book can’t pretend.

To maintain it concrete, we’ll use one actual venture the entire method via: the DoorDash Supply Length Prediction knowledge venture. It is a free venture, so you may comply with alongside and construct this your self.

We’ll work via it inside StrataScratch’s built-in pocket book atmosphere, an built-in Marimo pocket book you may open by clicking “Begin Fixing” on the venture web page, so there’s nothing to put in earlier than you begin.

So this is what we’ll do. We’ll take that one venture and run it via the 9 levels of an actual knowledge science venture: framing the enterprise drawback, pulling the info with SQL, cleansing it in Python, exploring it, engineering options, constructing and evaluating fashions, and at last deploying the outcome as an API and a dashboard that ends with a suggestion.

Every stage is a chapter of the identical story, and each is one thing a hiring supervisor can see for themselves. By the top, you may have a template you may drop nearly any venture into.

 
Building End to End Data Science Portfolio Project
 

# Beginning With a Enterprise Downside

 
Earlier than any code, determine what you are truly fixing.

The DoorDash venture offers us a clear enterprise query: given an order, how lengthy will supply take? That framing issues. It is about what the enterprise cares about, not the algorithm.

That is the primary place most portfolios go fallacious.

A venture titled “Supply Time Prediction” tells a hiring supervisor what you probably did for the corporate. A venture titled “XGBoost Regression Demo” tells them you adopted a tutorial.

Body the issue across the end result, and decide one thing with actual stakes: churn, forecasting, fraud, or, in our case, operational effectivity.

 

# Extracting the Knowledge With SQL

 
The DoorDash venture fingers us a CSV, historical_data.csv:

 
Building End to End Data Science Portfolio Project
 

However that isn’t the place knowledge lives in the actual world. In an organization, this dataset would come out of a database, and you would be the one writing the SQL to construct it.

We simulate this through the use of the built-in pocket book talked about earlier, because the dataset is already imported (as df). We straight question it with SQL. (If it had been an precise database, you’d question it with FROM historical_data.)

SELECT
    market_id,
    created_at,
    actual_delivery_time,
    store_id,
    store_primary_category,
    order_protocol,
    total_items,
    subtotal,
    total_onshift_dashers,
    total_busy_dashers,
    total_outstanding_orders
FROM df
WHERE actual_delivery_time IS NOT NULL
  AND actual_delivery_time > created_at;

 

Outputs:

market_id created_at actual_delivery_time … total_outstanding_orders
1 2015-02-06 22:24:17 2015-02-06 23:27:16 … 21
2 2015-02-10 21:49:25 2015-02-10 22:56:29 … 2
3 2015-01-22 20:39:28 2015-01-22 21:09:09 … 0
3 2015-02-03 21:21:45 2015-02-03 22:13:00 … 2
3 2015-02-15 02:40:36 2015-02-15 03:20:26 … 9
… … … … …
1 2015-02-08 19:24:33 2015-02-08 20:01:41 … 23

 

That’s value displaying in your portfolio.

As an alternative of quietly loading a file, describe the question that may produce your dataset: the joins throughout order, dasher, and retailer tables, the WHERE filters that drop unhealthy rows, and the GROUP BY clauses that do the heavy filtering and becoming a member of in SQL — and pull an analysis-ready desk into Python, not a uncooked dump.

 

# Cleansing the Knowledge in Python

 
Now we carry the info into Python. That is the unglamorous stage that’s 60 to 80 % of actual knowledge science work, and skipping it is without doubt one of the clearest alerts of inexperience.

For the DoorDash knowledge, cleansing means computing our goal (precise supply length is the supply timestamp minus the order creation timestamp), fixing varieties, and dealing with lacking and unimaginable values.

We use pandas for this, which is the precise default at portfolio scale.

df["created_at"] = pd.to_datetime(df["created_at"])
df["actual_delivery_time"] = pd.to_datetime(df["actual_delivery_time"])

# Our goal: how lengthy the supply truly took, in seconds
df["delivery_duration_seconds"] = (
    df["actual_delivery_time"] - df["created_at"]
).dt.total_seconds()

# Drop lacking and unimaginable values
# An actual supply is normally between 6 minutes and some hours
df2 = df[df["delivery_duration_seconds"].between(60, 3 * 3600)]

df3 = df2.dropna(subset=["delivery_duration_seconds"])

df3

 

Outputs:

market_id created_at actual_delivery_time delivery_duration_seconds
1 2015-02-06 22:24:17 2015-02-06 23:27:16 3779.0
2 2015-02-10 21:49:25 2015-02-10 22:56:29 4024.0
3 2015-01-22 20:39:28 2015-01-22 21:09:09 1781.0
3 2015-02-03 21:21:45 2015-02-03 22:13:00 3075.0
… … … …
3 2015-02-15 02:40:36 2015-02-15 03:20:26 2390.0

 

In case your dataset had been giant sufficient to pressure reminiscence, Polars can be the sooner, multi-core various, however for a venture like this, pandas is lots.

import polars as pl

df = pl.read_csv(
    "historical_data.csv",
    null_values=["NA"],
    try_parse_dates=True
)

df = df.with_columns(
    (pl.col("actual_delivery_time") - pl.col("created_at"))
    .dt.total_seconds()
    .alias("delivery_duration_seconds")
).filter(pl.col("delivery_duration_seconds") > 0)

 

# Exploring the Knowledge

 
Exploratory knowledge evaluation (EDA) is the place we discover the story we’ll finally inform.

The workflow is straightforward and repeatable: summarize the info with strategies like df.data() and df.describe(), then visualize distributions and relationships, then notice what’s stunning.

df3["delivery_minutes"] = df3["delivery_duration_seconds"] / 60
df3["delivery_minutes"].describe()

 

Notice that df3 is the cleaned dataset from the earlier pandas code.

Outputs:

statistic worth
depend 197283.0
imply 47.5
std 18.0
min 1.7
25% 35.1
50% 44.3
75% 56.3
max 179.8

 

For supply length, we would take a look at the way it varies by market, by hour of day, and by how busy the dashers are. We use Matplotlib and Seaborn for histograms, boxplots, and scatter plots.

import matplotlib.pyplot as plt
import seaborn as sns

# Distribution of supply time
sns.histplot(df3["delivery_minutes"].clip(higher=120), bins=50)

plt.xlabel("Supply length (minutes)")

 

Outputs:

 
Building End to End Data Science Portfolio Project
 

# the way it varies throughout markets
df3.groupby("market_id")["delivery_minutes"].median().sort_values()

 

Outputs:

market_id worth
1 46.9
2 43.3
5 43.4
6 43.6
3 44.1
4 44.4

 

The objective is to grasp what drives the factor you are predicting.

 

# Engineering the Options

 
Uncooked columns hardly ever make one of the best predictors. Function engineering is the place area considering turns into mannequin inputs, and it is typically what separates venture from a forgettable one.

Within the DoorDash venture, that is essentially the most fascinating stage. We construct a busy_dashers_ratio to seize how stretched the fleet is, and an estimated_non_prep_duration that mixes driving and order-placement time.

import numpy as np

df3["busy_dashers_ratio"] = (
    df3["total_busy_dashers"]
    / df3["total_onshift_dashers"]
)

df3["estimated_non_prep_duration"] = (
    df3["estimated_store_to_consumer_driving_duration"]
    + df3["estimated_order_place_duration"]
)

# The busy ratio can divide by zero
df3 = df3.substitute([np.inf, -np.inf], np.nan)

df3[
    [
        "busy_dashers_ratio",
        "estimated_non_prep_duration",
    ]
].head()

 

index busy_dashers_ratio estimated_non_prep_duration
0 0.424242 1307.0
1 2.000000 1136.0
2 0.000000 1136.0
3 1.000000 735.0
4 1.000000 1096.0

 

We flip categorical columns like market and order protocol into dummy variables. Then we cope with options that carry the identical data, utilizing a correlation heatmap and Variance Inflation Issue (VIF) to drop the redundant ones.

from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.pipeline import Pipeline

numeric = [
    "busy_dashers_ratio",
    "estimated_non_prep_duration",
    "total_items",
    "subtotal",
    "num_distinct_items",
    "min_item_price",
    "max_item_price",
    "total_onshift_dashers",
    "total_outstanding_orders",
]

categorical = ["market_id", "order_protocol"]

preprocess = ColumnTransformer([
    ("num", StandardScaler(), numeric),
    ("cat", OneHotEncoder(handle_unknown="ignore"), categorical),
])

# 11 uncooked columns grow to be 22 model-ready options after encoding
preprocess.fit_transform(df3[numeric + categorical].dropna()).form

 

Output:

 

Wrap all of this in a scikit-learn pipeline so the identical steps run identically on coaching and new knowledge, which quietly prevents knowledge leakage.

 

# Constructing the Mannequin

 
Resist the urge to leap straight to a elaborate mannequin.

Begin with a baseline, even a naive one which predicts the typical supply time. In case your actual mannequin cannot beat that, one thing is fallacious, and also you need to know early.

From there, we strive progressively stronger fashions: linear fashions like Ridge, then tree-based fashions, and gradient boosting with XGBoost.

from sklearn.model_selection import train_test_split
from sklearn.dummy import DummyRegressor
from sklearn.linear_model import Ridge
from sklearn.metrics import mean_squared_error
from xgboost import XGBRegressor

knowledge = df3[numeric + categorical + ["delivery_duration_seconds"]].dropna()

X = knowledge[numeric + categorical]
y = knowledge["delivery_duration_seconds"]

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.2,
    random_state=42
)

fashions = {
    "Baseline (imply)": DummyRegressor(technique="imply"),
    "Ridge": Ridge(),
    "XGBoost": XGBRegressor(
        n_estimators=600,
        learning_rate=0.05,
        max_depth=7,
        subsample=0.8,
        colsample_bytree=0.8,
        random_state=42
    ),
}

for identify, mannequin in fashions.objects():
    pipe = Pipeline([("pre", preprocess), ("model", model)])
    pipe.match(X_train, y_train)

    rmse = mean_squared_error(y_test, pipe.predict(X_test)) ** 0.5

    print(f"{identify}: RMSE = {rmse:.0f} sec")

 

Output:

mannequin RMSE
Baseline (imply) 1074 sec
Ridge 927 sec
XGBoost 875 sec

 

Tree-based fashions normally carry out finest on tabular enterprise knowledge like this. In your writeup, clarify why you selected what you selected.

That reasoning is what a hiring supervisor reads to see whether or not you perceive the instruments or simply imported them.

 

# Evaluating Actually

 
A single accuracy quantity proves nothing. For a regression drawback like supply length, we report an error metric resembling root imply squared error (RMSE) and examine each mannequin towards our baseline and towards one another.

The larger level is validating actually. Use cross-validation as a substitute of trusting one fortunate train-test break up, and by no means tune your mannequin towards the take a look at set, as a result of the second you do, your reported rating turns into optimistic fiction.

from sklearn.model_selection import cross_val_score

pipe = Pipeline([("pre", preprocess), ("model", models["XGBoost"])])

scores = cross_val_score(
    pipe,
    X,
    y,
    cv=5,
    scoring="neg_root_mean_squared_error"
)

print("Fold RMSEs:", (-scores).spherical().astype(int))
print(f"CV RMSE: {-scores.imply():.0f} sec (+/- {scores.std():.0f})")

 

Output:

metric worth
Fold RMSEs [900, 886, 867, 878, 882]
CV RMSE 883 sec ±11 sec

 

For classification issues, report precision, recall, and F1 alongside accuracy, not accuracy alone.

 

# Deploying the Mannequin

 
Right here is the place most portfolios merely cease, which is precisely why going additional makes yours stand out. Wrapping the mannequin in an API is what lets anybody truly use it.

We serialize the educated mannequin with joblib, then wrap it in a small service utilizing FastAPI, which supplies us request validation and computerized docs with nearly no effort.

import joblib

pipe.match(X_train, y_train)
joblib.dump(pipe, "delivery_model.joblib")

# api.py
from fastapi import FastAPI
from pydantic import BaseModel
import joblib
import pandas as pd

app = FastAPI()
mannequin = joblib.load("delivery_model.joblib")


class Order(BaseModel):
    busy_dashers_ratio: float
    estimated_non_prep_duration: float
    total_items: int
    subtotal: float
    num_distinct_items: int
    min_item_price: float
    max_item_price: float
    total_onshift_dashers: float
    total_outstanding_orders: float
    market_id: int
    order_protocol: int


@app.publish("/predict")
def predict(order: Order):
    row = pd.DataFrame([order.model_dump()])
    seconds = float(mannequin.predict(row)[0])
    return {"predicted_delivery_seconds": spherical(seconds)}

 

A POST to /predict now returns one thing like {"predicted_delivery_seconds": 2472}. We bundle all the things in a Docker container and deploy it someplace public, like a free cloud tier. Now anybody can ship an order and get a predicted supply time again.

 

# Constructing a Dashboard

 
The ultimate stage closes the loop again to stage one. Not everybody reviewing your work will name your API, so give them one thing to click on. We construct a small dashboard with Streamlit, the quickest solution to flip a Python script into an interactive app.

For our venture, the dashboard lets somebody enter order particulars and see the expected supply time, and discover which components push it up or down.

import streamlit as st
import joblib
import pandas as pd

mannequin = joblib.load("delivery_model.joblib")

st.title("Supply Length Predictor")

order = {
    "busy_dashers_ratio": st.slider(
        "Busy dashers ratio",
        0.0,
        2.0,
        0.5
    ),
    "estimated_non_prep_duration": st.number_input(
        "Non-prep length (sec)",
        worth=900
    ),
    "total_items": st.number_input(
        "Whole objects",
        worth=4,
        step=1
    ),
    "subtotal": st.number_input(
        "Subtotal (cents)",
        worth=3441
    ),
    "num_distinct_items": st.number_input(
        "Distinct objects",
        worth=4,
        step=1
    ),
    "min_item_price": st.number_input(
        "Min merchandise worth",
        worth=557
    ),
    "max_item_price": st.number_input(
        "Max merchandise worth",
        worth=1239
    ),
    "total_onshift_dashers": st.number_input(
        "On-shift dashers",
        worth=33
    ),
    "total_outstanding_orders": st.number_input(
        "Excellent orders",
        worth=21
    ),
    "market_id": st.selectbox(
        "Market",
        [1, 2, 3, 4, 5, 6]
    ),
    "order_protocol": st.selectbox(
        "Order protocol",
        [1, 2, 3, 4, 5, 6, 7]
    ),
}

if st.button("Predict"):
    seconds = float(
        mannequin.predict(
            pd.DataFrame([order])
        )[0]
    )

    st.metric(
        "Predicted supply time",
        f"{seconds / 60:.1f} min"
    )

 

Then we finish the place good knowledge science initiatives all the time finish: with a suggestion. If a excessive busy_dashers_ratio is the largest driver of lengthy deliveries, the enterprise motion is to regulate staffing throughout peak load. Finish with the enterprise motion, not simply the prediction.

 

# Conclusion

 
The lifecycle is the differentiator. Anybody can prepare a mannequin, however only a few candidates carry an issue all the best way from a SQL question to a deployed app with a transparent suggestion on the finish. That end-to-end story is what a portfolio is for, and it is what a resume can by no means present by itself.

 
Building End to End Data Science Portfolio Project
 

It helps to see what every half of that story proves. The early levels — framing the issue, writing the SQL, cleansing and exploring the info — present that you could flip a messy enterprise query into one thing a mannequin can truly be taught from. The later levels — evaluating actually, deploying, and constructing a dashboard — present that you could take a mannequin out of a pocket book and put it in entrance of somebody who has to decide.

Most candidates can do one half. Those who do each are uncommon, and that’s precisely the hole you are closing. The through-line that ties it collectively is the enterprise framing from stage one: each stage ought to level again on the query you began with and ahead to the advice you finish on.

You do not have to invent a venture to follow this. The DoorDash venture we used right here is one in every of many actual firm take-homes in StrataScratch’s knowledge initiatives, alongside issues from Meta, Capital One, Google, and others.

Choose one, run it via all 9 levels, and write it up actually, together with the elements that did not work. That single completed venture will do extra in your job search than 5 extra notebooks that cease on the mannequin.
 
 

Nate Rosidi is a knowledge scientist and in product technique. He is additionally an adjunct professor instructing analytics, and is the founding father of StrataScratch, a platform serving to knowledge scientists put together for his or her interviews with actual interview questions from prime corporations. Nate writes on the most recent traits within the profession market, offers interview recommendation, shares knowledge science initiatives, and covers all the things SQL.



READ ALSO

Buyer Expertise Analytics: What Model Knowledge Reveals

Contained in the FTC’s Information Breach Playbook for 2026 |

Tags: BuildingDataEndtoEndPortfolioprojectScience

Related Posts

Customer experience analytics what brand data reveals featured.jpg
Data Science

Buyer Expertise Analytics: What Model Knowledge Reveals

August 12, 2026
Ftc data breach notification matrix 2026.png
Data Science

Contained in the FTC’s Information Breach Playbook for 2026 |

August 12, 2026
Kdn the ultimate guide to contributing to open source projects feature.png
Data Science

The Final Information to Contributing to Open Supply Tasks

August 11, 2026
Ai drug discovery companies leading development platforms featured.png
Data Science

AI Drug Discovery Corporations: Main Improvement Platforms

August 11, 2026
Qwen 38 max alibaba ai pricing datacenter.jpg.png
Data Science

Qwen3.8-Max Alerts Alibaba’s Wager That Low-cost Beats Sensible |

August 10, 2026
Kdn chugani minimal ai engineer toolkit 2026 feature.png
Data Science

The Minimal AI Engineer Toolkit for 2026

August 10, 2026

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

POPULAR NEWS

Gemini 2.0 Fash Vs Gpt 4o.webp.webp

Gemini 2.0 Flash vs GPT 4o: Which is Higher?

January 19, 2025
Chainlink Link And Cardano Ada Dominate The Crypto Coin Development Chart.jpg

Chainlink’s Run to $20 Beneficial properties Steam Amid LINK Taking the Helm because the High Creating DeFi Challenge ⋆ ZyCrypto

May 17, 2025
Image 100 1024x683.png

Easy methods to Use LLMs for Highly effective Computerized Evaluations

August 13, 2025
Blog.png

XMN is accessible for buying and selling!

October 10, 2025
0 3.png

College endowments be a part of crypto rush, boosting meme cash like Meme Index

February 10, 2025

EDITOR'S PICK

Gemini generated image 1rsfbq1rsfbq1rsf scaled 1.jpg

Cease Treating AI Reminiscence Like a Search Downside

April 12, 2026
MulticollinearityPhoto.jpg

Why Your Betas Explode: The Hidden Geometry of Multicollinearity

July 16, 2026
3440604A B881 4555 8517 C8E5FA9743C3.jpg

MCP Defined: How Fashionable AI Brokers Connect with the Actual World

July 29, 2026
Mined in america act.jpeg

US Senators Push ‘Mined in America Act’ to Safe Bitcoin Mining and Reserve

April 1, 2026

About Us

Welcome to News AI World, your go-to source for the latest in artificial intelligence news and developments. Our mission is to deliver comprehensive and insightful coverage of the rapidly evolving AI landscape, keeping you informed about breakthroughs, trends, and the transformative impact of AI technologies across industries.

Categories

  • Artificial Intelligence
  • ChatGPT
  • Crypto Coins
  • Data Science
  • Machine Learning

Recent Posts

  • Constructing an Finish-to-Finish Knowledge Science Portfolio Mission
  • Coinbase Expands Derivatives Buying and selling To UK Skilled Shoppers
  • Utilizing a Transformer Mannequin: From Coaching to Inference
  • Home
  • About Us
  • Contact Us
  • Disclaimer
  • Privacy Policy

© 2024 Newsaiworld.com. All rights reserved.

No Result
View All Result
  • Home
  • Artificial Intelligence
  • ChatGPT
  • Data Science
  • Machine Learning
  • Crypto Coins
  • Contact Us

© 2024 Newsaiworld.com. All rights reserved.

Are you sure want to unlock this post?
Unlock left : 0
Are you sure want to cancel subscription?