AI doesn’t just write code faster—it puts junior engineers into the captain’s chair before they know how to read the radar.
Suddenly, we can do more in less time. What used to take weeks or months can take hours to be done nowadays using AI. But the implications of that sound more complex than assuming that anyone can release successful products today. The recent changes in how things are done in software engineering generally impact its specialized branch, data engineering, too.
An expansion of the horizon is the first description that comes to mind to show how things changed. It’s like someone who’d been able to do an engineering role in a ship before, but suddenly having the capacity to run the whole ship. Unlike the view that we can rest and let AI do things, in reality our vision has to expand now. If you weren’t the captain and didn’t need to look at the radar, now you have to.
Many data engineers’ work is primarily around data warehouse solutions. That includes the work to extract, transform, and load data into warehouses (ETL), or managing the data warehouse solutions. Obviously SQL is one of the most replaceable skills, along with the jobs to automate, orchestrate, and clean the data. But the knowledge of this type of engineers in their organizations’ data isn’t. Before AI, enterprise search faced serious challenges while global search engines were thriving. I don’t see any difference here. While enterprise search required crawlers to be allowed everywhere in the organization considering the different security permissions, AI agents would require something similar to be enabled, but for data lineage. There are many reasons to question how realistic it is for AI to govern data in order to do this type of engineering, but not all AI involvement is about “replacing” us, especially what I’m pointing out here.
Getting into the details of the data engineer role, the gap reduces between those who work on data warehouses, or those who work on real-time and batch pipelines for all other purposes. This area is more of software engineering in a wider sense. This includes: designing and implementing microservices and APIs, Spark and Airflow scripts, in addition to the wide knowledge of data storage, data streaming, data governance, and data analysis tools. Agentic AI impacted these mainly through how quickly code can be written and how quickly a solution could be deployed. But here, we’re talking about less than a half of the software product lifecycle components which we’re familiar with in software engineering.
Back to the vision, and the expansion in horizon, where does it matter in the software life cycle? Coding is the first place that comes to us, but as much as it seems revolutionary and can replace humans, in fact the possibility of disaster is higher now. In the example of the ship, at least the bad sailor would never make it to be a captain, but now they can build ships and go to the sea until they fail. Unless they had a good vision.
The challenge of designing data engineering solutions in the age of AI
In a dictatorship, a powerful man entered a state factory and asked them to double the production. The technicians said that’s impossible because of the horsepower of the machines they have. The powerful man said: “I want you to go to the market tomorrow and to buy all the horsepower in the market.” Surely it was difficult to correct this guy. Somehow, anyone of us can be this man and the AI can end up doing what we want even if it lacked common sense or wisdom.
Let’s take the example of designing a data engineering solution for moving data from point A and store it in B. Storage solution can be a file system, object storage, relational database, documents stored in a document database, storing big data in a columnar database for analysis, or it could be data streaming. Retrieval can have different purposes too: it could be for high reads serving web viewers, archiving for less frequent retrieval, or many other possibilities. Now, AI can design a solution for you even if you just mentioned the primary description of moving data from A to B. But it can brainstorm and ask you further, yet, that would still be limited without the deep vision and knowledge of available solutions. And what if the AI brainstorming wasn’t enough? What if you didn’t have enough knowledge to dive into the brainstorming questions?
Let’s say that a brainstorming skill in Claude code had taken you deeper into the point of understanding that you need data streaming for the solution above rather than to store data (even if you expressed it as data storage). You reached a point that the data is only required for a limited period to be accessed with high-speed reads by a microservice, then it should be deleted. Let’s say that the AI along with the brainstorming questions have led you to the fact that you need a solution like SQS or Kafka, for data streaming. But is there any guarantee that the AI “skills” questions can go further to something like topics design and how the records should be stored in terms of keys and partitioning? AI can come with a default option and implement that, without the knowledge about the future of the solution usage and the data quantity.
We need to cover more concepts than any other time before. Not that we, as software engineers or as data engineers, were just lazy (most of us are), but it wasn’t realistic to learn 300 concepts that you’ll forget while your work is concentrated around 30 ones. Today, if 80% of what used to take your time has been shifted to AI, you’re suddenly in a lead position even if your title doesn’t say so. You’re automatically in a captain position and have to know more about other parts of the ship, and can learn too.
Back to the previous example, what happened to me in a similar situation with one of the great LLM models, it didn’t even reach the point to suggest data streaming. The advanced model looked at my limited requests assuming that’s how it’s always going to be, and it just suggested a table in a relational database for the data to reside temporarily as a message queue. What if my requests were in thousands per second and I needed data streaming rather than a relational database? We didn’t even reach the point to discuss keys, topics, or partitioning that I assumed in the previous example.
The reason why my agentic AI design decision went into that direction is that it’s mostly trained on coding and solutions done by amateurs, and open source solutions, but not large enterprise solutions designed by senior engineers. You expect your normal text to be generated with training data of the gen AI model coming from the language of the Guardian, the BBC or Shakespeare, but you can never generate your code based on logic and algorithms written by Google developers.
If engineering means designing a solution with the available tools to solve a problem, the horizon of the problem that AI sees is completely different from what you see. Vision is everything today, and vision doesn’t only mean what you want to achieve, but how it’s going to develop over time, how debugging is going to work, and how performance tuning is going to develop over time. If you had the vision to cover all of these, then AI can achieve them for you, but having a blind spot yourself, then you may end up with cliché AI suggestions like Postgres and Supabase that Claude really likes (for what AI is trained on).
Is any of that new? Not to me. We always designed solutions like that, but while implementation was a big deal before, we were sure at least that whoever is designing and leading a small team to implement it, they’d be knowledgeable of what’s going on. Today, anyone can do that with millions of tokens wasted, to make the solution capable of standing, unable to walk, while it’s expected to run. The waste of tokens, testers’ time, and deployment/post-deployment disaster is more possible today.
What makes this dangerous is the illusion of competence: you think that because AI produces good syntax and working code, naturally it is good at design. No need to say how little we find of design use cases when we search online; most of what exists is produced by cloud solution providers to promote their products. If AI is good at coding, or writing essays, that doesn’t make it a good designer.
Another example is using AI to tune EKS. Let’s say you have weekly monitoring data, EKS applications resources consumption metrics, and you provided these to AI to help you find the right metrics for scaling and properly allocate the resources you need. Sounds perfect? Well, again, very advanced models didn’t even ask me about monthly or annual CPU and memory usage metrics. Imagine taking the values recommended by AI, and heading towards the nearest monthly peak.
So while the tool is as powerful as a magic wand, the outcome is still limited by your vision of the problem. And while we were more relaxed regarding the vision, as the time and effort required would keep us focused on a limited scope, there’s more pressure now to look in all possible directions and to ask all the possible questions, such as anticipating the time range of CPU and memory peaks in the example above, which isn’t possible without wider knowledge of both the behavior of our systems and the engineering concepts behind them. You might be still working in the ship, and not responsible for the final sailing decisions, but you can clearly see now the whole picture, better than any time before.
