• Home
  • About Us
  • Contact Us
  • Disclaimer
  • Privacy Policy
Monday, August 10, 2026
newsaiworld
  • Home
  • Artificial Intelligence
  • ChatGPT
  • Data Science
  • Machine Learning
  • Crypto Coins
  • Contact Us
No Result
View All Result
  • Home
  • Artificial Intelligence
  • ChatGPT
  • Data Science
  • Machine Learning
  • Crypto Coins
  • Contact Us
No Result
View All Result
Morning News
No Result
View All Result
Home Artificial Intelligence

I Thought Loading Information Was the End Line. It Was the Beginning Level.

Admin by Admin
August 9, 2026
in Artificial Intelligence
0
RSS transform article dbt.jpg
0
SHARES
0
VIEWS
Share on FacebookShare on Twitter

READ ALSO

The Drawback with pandas Isn’t Efficiency. It’s Cognitive Overhead.

Constructing a Streamlit UI for My LangGraph AI Agent


, I gave myself a 12-month roadmap to go from information analyst to information engineer. I’m solely about two months into it. In that quick stretch I’ve already constructed two ETL pipelines from scratch, the primary one pulling GitHub repo information into SQLite, the second pulling RSS articles into PostgreSQL with Docker and Kestra dealing with the orchestration. I wrote about scheduling that second pipeline to run robotically each hour, and on the time, that felt like an actual milestone. The info was flowing in by itself, no guide runs, no me remembering to set off something.

However someplace between writing that article and beginning this one, I ran a question by myself information and realized one thing. I couldn’t type my articles by date correctly. I couldn’t inform which blogs have been publishing probably the most. The info had been sitting in Postgres for weeks, technically “loaded,” and I hadn’t really checked out it intently till I wanted it for one thing.

Seems I’d constructed two pipelines and skipped the half that makes the information helpful. Extract, load, after which nothing. No transformation, no modeling, no actual construction previous “it’s in a desk now.”

This text is about fixing that. I lastly sat down and discovered dbt, and within the course of discovered what “evaluation prepared” really means, as a result of it seems loading information and having usable information are two very various things.

The Information Was Loaded. It Simply Wasn’t Usable.

Right here’s what my articles desk really seemed like as soon as I finished and paid consideration to it.

The schema itself was easy, actually about so simple as a desk can get:

CREATE TABLE IF NOT EXISTS articles (
    id TEXT PRIMARY KEY,
    title TEXT NOT NULL,
    hyperlink TEXT NOT NULL,
    abstract TEXT,
    revealed TEXT
);

Discover that final column. revealed is a TEXT area. Not a timestamp, not a date, only a plain string that occurred to seem like a date in the event you squinted at it.

Once I queried the ten most up-to-date articles, that is what got here again:

title                                                    | revealed
----------------------------------------------------------+---------------------------------
Django Weblog: Final Name 2026 Django Developer Survey      | Wed, 08 Jul 2026 19:31:21 +0000
Mike Driscoll: New E-book Launch: Python Typing              | Wed, 08 Jul 2026 18:46:18 +0000

That appears advantageous at a look. It’s readable. However attempt to really do something with it. Need the articles from the final 7 days? You’ll be able to’t filter on that with out casting it first, each single time, in each single question. Wish to type chronologically and belief the order? Textual content sorting and date sorting aren’t the identical factor, and relying on the format, they’ll quietly disagree with one another.

Then there was the second drawback, the one I nearly missed totally as a result of it was hiding in plain sight. Take a look at these titles once more:

Django Weblog: Final Name 2026 Django Developer Survey
Mike Driscoll: New E-book Launch: Python Typing

Each single title on this feed follows the identical sample. Writer or weblog identify, a colon, then the precise headline. That’s actual, structured info sitting inside a single textual content area, utterly unusable as a filter or a group-by. I couldn’t reply a query so simple as “which blogs submit probably the most on Planet Python” as a result of that info wasn’t a column. It was simply textual content, buried.

In order that was the precise state of issues. Two pipelines constructed, information flowing in on schedule, and I nonetheless couldn’t reply primary questions on my very own information. Loading it was by no means the end line. I simply hadn’t gotten to the place to begin but.

Why dbt, Particularly

My first intuition was to simply repair this in Python, since that’s the device I already belief. Write a script that reads from articles, parses the dates, splits the titles, writes the outcomes into new columns or a brand new desk. And that might have labored, technically.

However the extra I thought of it, the extra that felt like patching the identical gap I’d already dug twice. Each of my pipelines have been extract and cargo, full cease, and if I bolted transformation logic onto a Python script once more, I’d simply be including a 3rd untested, undocumented step to a system that already had two. I wouldn’t be studying something new. I’d simply be writing extra of the identical factor I already knew find out how to write.

dbt does this in another way, and that distinction is sort of the entire level of the device. As an alternative of a script that runs as soon as and produces some output you must belief blindly, dbt fashions are SQL that will get model managed, examined, and documented as a part of the identical workflow. You write a change, and in the identical undertaking you’ll be able to assert issues about it: this column ought to by no means be null, this ID ought to all the time be distinctive. If these assumptions break, you discover out instantly, not three weeks later when a chart seems to be flawed and you haven’t any concept why.

It additionally matches how the trade really works. Each information engineering job submit I’ve checked out over the previous two months mentions dbt, or one thing dbt-shaped. Studying it wasn’t nearly fixing my RSS information, it was about studying the device that’s develop into the default manner groups deal with the “T” in ETL.

So as a substitute of one other Python script, I made a decision to truly sit down and be taught dbt correctly, on information I already had, with issues I already understood. Right here’s how that went.

Setting Up (and Instantly Hitting a Wall)

Getting dbt put in ought to have been the boring half. It wasn’t.

I attempted pip set up dbt-postgres and received a wall of dependency decision errors, dbt-core had no matching distribution for my setting. Seems I used to be operating Python 3.14, which is new sufficient that dbt hadn’t caught as much as it but. dbt Core formally helps as much as 3.13 proper now, and there’s often a lag earlier than it helps no matter Python simply launched.

The repair wasn’t sophisticated as soon as I understood the precise drawback, set up an older, supported Python model alongside my current one, and construct a digital setting particularly for dbt utilizing that:

py -3.12 -m venv dbt-env
dbt-envScriptsactivate
pip set up dbt-postgres

That’s a small factor, but it surely’s the sort of small factor that eats an hour in the event you don’t know to search for it, and I feel it’s value together with right here as a result of it’s precisely the sort of setup friction that doesn’t present up in tutorials. Tutorials assume your setting already works. Mine didn’t, and I’d guess I’m not the one one operating a Python model that’s forward of what dbt at the moment helps.

As soon as that was sorted, connecting dbt to my current Postgres database (already operating regionally in Docker from my RSS pipeline) was simple. dbt init, decide postgres, plug within the host, port, credentials, and database identify, and dbt debug confirms the connection:

All checks handed!

With that, I had a dbt undertaking sitting on prime of the identical Postgres occasion my RSS pipeline had been writing to for weeks. Time to truly repair the information.

Constructing the Staging Mannequin

The primary actual dbt idea I needed to perceive was the distinction between a supply and a mannequin. My uncooked articles desk isn’t one thing dbt constructed, it’s exterior information that already exists, so dbt calls it a supply. I outlined that in a sources.yml file, which is actually simply dbt’s manner of formally acknowledging “this desk exists, and I depend upon it”:

sources:
  - identify: rss_pipeline
    schema: public
    tables:
      - identify: articles

From there, I constructed my very first mannequin, stg_articles, a staging mannequin whose total job is to wash up the uncooked information with out doing something fancy but. That is the place each of my unique issues received fastened in the identical file.

For the date:

to_timestamp(revealed, 'Dy, DD Mon YYYY HH24:MI:SS OF') as published_at

For the buried writer identify:

split_part(title, ':', 1) as writer,
trim(substring(title from place(':' in title) + 1)) as article_title

I ran dbt run, then went and truly queried the consequence as a substitute of assuming it labored:

published_raw                     published_at
Solar, 05 Jul 2026 16:29:47 +0000   2026-07-05 16:29:47+00

Actual timestamps. And after I checked the writer break up:

writer                      | article_title
Python Software program Basis  | Python Packaging Council Inaugural Election Dates

Clear separation, even on titles with multiple colon in them, like a PyCoder’s Weekly problem title that had a colon in each the supply identify and the headline itself. The break up logic solely breaks on the primary colon, so it held up advantageous.

Including Assessments

That is the half that made the entire undertaking really feel much less like “I wrote some SQL” and extra like precise engineering. I added exams straight in a schema file subsequent to the mannequin:

columns:
  - identify: article_id
    exams:
      - distinctive
      - not_null
  - identify: published_at
    exams:
      - not_null

Operating dbt check doesn’t simply verify that the SQL runs, it checks that my assumptions concerning the information really maintain:

PASS=4 WARN=0 ERROR=0 SKIP=0 NO-OP=0 REUSED=0 TOTAL=4

That not_null check on published_at particularly is the one that might have caught it if my date format string had been flawed. As an alternative of silently producing nulls I may not discover for weeks, I’d have seen a failed check the second I ran it.

Constructing a Mart, and Lastly Asking a Actual Query

With clear staging information in place, I constructed yet one more mannequin on prime of it, articles_by_author, which aggregates the information into one thing I may really ask a query of: which blogs submit probably the most, and the way not too long ago.

choose
    writer,
    rely(*) as total_articles,
    max(published_at) as most_recent_article,
    min(published_at) as earliest_article
from {{ ref('stg_articles') }}
group by writer
order by total_articles desc

That ref() perform as a substitute of supply() issues right here, it’s how dbt is aware of this mannequin is determined by stg_articles, not on the uncooked desk straight. That dependency monitoring is what builds the lineage graph later.

The consequence was the primary genuinely new factor I may see on this information since I began amassing it two months in the past:

writer                       total_articles   most_recent_article
Python Software program Basis   5                2026-07-09 14:11:06+00
Django Weblog                4                2026-07-08 19:31:21+00

A query I couldn’t reply per week earlier, answered in a single question, on information I’d already had sitting round the entire time.

Seeing the Complete Factor

The final step was operating dbt docs generate and dbt docs serve, which builds an interactive documentation website with a lineage graph, principally a visible map of how information flows via the undertaking. Mine confirmed precisely three linked nodes:

The place This Leaves Me

In order that’s the undertaking. Two clear fashions, seven passing exams, and a lineage graph that truly exhibits an actual chain from uncooked information to one thing I can ask questions of. In comparison with the place I began this piece, unable to type by date or inform which blogs posted probably the most, that’s an actual shift, even when the underlying dataset didn’t change in any respect. Identical information. Very completely different usefulness.

I wish to be trustworthy about what this isn’t, although. That is nonetheless operating totally by myself machine. The Postgres database, the dbt undertaking, all of it lives regionally in Docker, which suggests none of this exists wherever the second my laptop computer is off. There’s additionally just one RSS feed feeding into this proper now, so “which blogs submit probably the most” is a reasonably small query with a reasonably small dataset behind it. And I haven’t touched something round alerting or monitoring if a check begins failing quietly within the background.

None of that takes away from what I really discovered right here, although. I feel there’s a distinction between a undertaking being completed and a undertaking having taught you what it was supposed to show you. This one did the second factor. I perceive the distinction between a supply and a mannequin now. I perceive why exams aren’t non-compulsory in the event you really wish to belief your personal information. And I perceive, in a really concrete manner this time, why “the information is loaded” and “the information is usable” are two utterly completely different claims.

The subsequent drawback is clear, actually. All the pieces I constructed right here nonetheless is determined by my laptop computer being on and Docker operating. That’s the following wall I’m going to hit, and doubtless the following factor I write about, taking this entire stack off my machine and placing it someplace it will probably really run with out me.

Two months right into a twelve month roadmap, and I feel that’s about proper. Slower than I’d like some days, however each wall I’ve hit to this point has taught me one thing I couldn’t have discovered by studying about it first.

Thanks for studying!

That is a part of my ongoing collection documenting my transition from techniques analyst to information engineer. If you happen to’ve been following alongside, thanks.

Join with me on LinkedIn, YouTube, and Twitter.

Tags: DataFinishLineLoadingPointStartingThought

Related Posts

Towards data science cognitive overhead.jpg
Artificial Intelligence

The Drawback with pandas Isn’t Efficiency. It’s Cognitive Overhead.

August 9, 2026
Towfiqu barbhuiya 9gPKrsbGmc unsplash scaled 1.jpg
Artificial Intelligence

Constructing a Streamlit UI for My LangGraph AI Agent

August 8, 2026
Pinning notes 8581050 v3 card.jpg
Artificial Intelligence

Loop Engineering for Itemizing Questions: When the Reply Is Each Passage, Not the Prime One

August 7, 2026
ChatGPT Image Aug 2 2026 10 19 34 PM.jpg
Artificial Intelligence

I Constructed an AI Knowledge Agent Which Can Question Knowledge and Reply Enterprise Questions. Right here’s How.

August 7, 2026
ChatGPT Image Jul 31 2026 09 05 45 PM.jpg
Artificial Intelligence

I Constructed a Instrument-Calling Agent in Python. Right here’s How I Debugged It

August 6, 2026
Fig a attention@3x scaled 1.jpg
Artificial Intelligence

How a Frontier Mannequin Will get Constructed, Learn from the Kimi K3 Report

August 5, 2026
Next Post
Bybit20Dubai20Office id 26e954e6 e167 4a36 919f 197d935b048f size900.jpg

Bybit Makes use of Tokenised Equities as Underlyings for Structured Yield

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

POPULAR NEWS

Gemini 2.0 Fash Vs Gpt 4o.webp.webp

Gemini 2.0 Flash vs GPT 4o: Which is Higher?

January 19, 2025
Chainlink Link And Cardano Ada Dominate The Crypto Coin Development Chart.jpg

Chainlink’s Run to $20 Beneficial properties Steam Amid LINK Taking the Helm because the High Creating DeFi Challenge ⋆ ZyCrypto

May 17, 2025
Image 100 1024x683.png

Easy methods to Use LLMs for Highly effective Computerized Evaluations

August 13, 2025
Blog.png

XMN is accessible for buying and selling!

October 10, 2025
0 3.png

College endowments be a part of crypto rush, boosting meme cash like Meme Index

February 10, 2025

EDITOR'S PICK

Rosidi the hidden curriculum of data science interviews 2.png

The Hidden Curriculum of Information Science Interviews: What Firms Actually Take a look at

October 23, 2025
Generative Ai Shutterstock 2273007347 Special.jpg

Report Findings – Safety Execs Determine GenAI because the Most Vital Threat for Organizations

September 28, 2024
0ey3ijiwnn5s Vdcv.jpeg

Depth-First Search — Elementary Graph Algorithm | by Robert Kwiatkowski | Sep, 2024

September 28, 2024
Phoenix id ebb0f7bb 05fd 4d90 810d b5fc25f2965f size900.jpeg

UAE Bitcoin Miner Phoenix Group Reviews 43% Income Decline, Highlights Shift Towards Digital Treasury

August 1, 2025

About Us

Welcome to News AI World, your go-to source for the latest in artificial intelligence news and developments. Our mission is to deliver comprehensive and insightful coverage of the rapidly evolving AI landscape, keeping you informed about breakthroughs, trends, and the transformative impact of AI technologies across industries.

Categories

  • Artificial Intelligence
  • ChatGPT
  • Crypto Coins
  • Data Science
  • Machine Learning

Recent Posts

  • Easy methods to Implement Structured Output with Native LLMs
  • Bybit Makes use of Tokenised Equities as Underlyings for Structured Yield
  • I Thought Loading Information Was the End Line. It Was the Beginning Level.
  • Home
  • About Us
  • Contact Us
  • Disclaimer
  • Privacy Policy

© 2024 Newsaiworld.com. All rights reserved.

No Result
View All Result
  • Home
  • Artificial Intelligence
  • ChatGPT
  • Data Science
  • Machine Learning
  • Crypto Coins
  • Contact Us

© 2024 Newsaiworld.com. All rights reserved.

Are you sure want to unlock this post?
Unlock left : 0
Are you sure want to cancel subscription?