Posts

  • data-analytics · ai 5 min read

    When a model earns its keep: anatomy of an ML project worth building

    The last post was about when not to build a model — the three tests a request has to survive before modelling is even on the table. This one is the other side of the ledger: what a model looks like when it genuinely clears the bar, and worth every hour of maintenance it will quietly demand for years. Because they do exist. The point was never “fewer models.” The point is models that earn their place, and those have a recognisable shape if you know what to look for.

  • data-analytics · ai 5 min read

    When NOT to build a model

    The fastest way to lose the plot on a data project is to start from the answer — “let’s build a model for that” — before anyone has said clearly what decision it’s supposed to improve. Machine learning has become the default ambition of far too many analytics requests, and the cost of that reflex is quietly enormous. Most problems asking for a model don’t need one. A well-built dashboard, a plain business rule, or even a genuinely cleaned dataset answers them better, faster, and in a way people actually trust.

  • data-analytics · practice 6 min read

    The dataset is never clean — a realistic first week on a messy dataset

    If you work with data long enough, you hear the same request in a dozen different clothes: “just clean the data and give me a dashboard.” It sounds like a two-line task. It never is. The honest reality is that the dataset on your desk is already dirty in ways you can’t yet see, and the people asking for it don’t know which of those ways actually matter for the decision they’re trying to make. So the first week isn’t about cleaning anything. It’s about finding out what kind of dirty you’re dealing with, and which of it is safe to ignore.

  • personal · introduction · data-ai 2 min read

    Hello, I'm Maximilian

    This is the first post on my new blog, so it seems only fair to start with a quick introduction.

  • data-engineering 16 min read

    Lakehouse Essentials: building a Databricks medallion stack as a one-person data team

    Building a lakehouse sounds like a daunting engineering task. It’s a phrase that summons images of a full platform team, months of infrastructure yak-shaving, and a budget line that makes finance wince. I wanted to write up the fundamental components I actually used to build a lakehouse on Databricks at GetGo — and along the way show that, for a single-person data operation, it’s genuinely more tractable than the marketing suggests. The reason is simple: Databricks bundles most of the hard parts out of the box, so one person can spend their time on data processing rather than infrastructure. You’ll soon see that it’s really easy. Let’s go.

subscribe via RSS