- Octoenergy
- Berlin, BE
- Full-Time
- 21 days ago
Data Engineer Mid-Level (m/w/d).
Before you go
Before you leave us, sign up for our email alerts
We don't do job spam, just the best digital jobs delivered straight to your inbox.
Data Engineer Mid-Level (m/w/d): our view in 3 lines...
- The Role:A mid-level data engineer role focused on building and operating a data and analytics platform for a German energy metering business.
- The Person:The person will build and tune lakehouse pipelines, model data, run data-quality and lineage checks, create BI and Streamlit data apps, and train and deploy ML use cases.
- Requirements:The role calls for 2–4 years of practice in data engineering, Databricks, dbt, Apache Airflow, Lightdash, Streamlit, SQL, Python, pandas, PySpark and scikit-learn.
About the role
Deine Aufgaben:
Lakehouse & Data Engineering (ca. 50%):
-
Konzeption, Aufbau und Tuning performanter Pipelines auf Basis von Databricks.
-
Datenmodellierung mit dbt sowie Orchestrierung robuster DAGs via Apache Airflow.
-
Governance & Performance:
-
Etablierung automatisierter Data-Quality-Prüfungen, Data Lineage und Cluster-Optimierung.
-
Analytics & Data Apps (ca. 25%):
-
Aufbau einer semantischen Metriken-Schicht und Self-Service-BI in Lightdash.
-
Entwicklung interaktiver Analyse-Tools und Data Apps in Python (Streamlit).
-
Data Science & ML (ca. 25%):
-
Explorative Analysen, Prototyping und Training von ML-Modellen (z. B. Prognosen, Anomalien).
-
MLOps & Business-Integration:
-
Experiment-Tracking auf Databricks und Bereitstellung der Ergebnisse für Fachbereiche.
-
-
-
-
Dein Profil:
Berufserfahrung:
-
Mindestens 2–4 Jahre Praxis im Data Engineering, Analytics Engineering oder einer vergleichbaren Rolle.
-
Core Engineering Stack:
-
Fundierte Praxis mit Databricks (Delta Lake, Unity Catalog), dbt und Apache Airflow.
-
BI & Data Apps:
-
Erfahrung mit Lightdash (oder Interesse an code-basierter BI) sowie mit Streamlit.
-
Code & Datenbanksprachen:
-
Exzellentes SQL (Performance-Tuning, komplexe Joins) und sehr gutes Python (pandas, PySpark, scikit-learn).
-
Methodik & Best Practices:
-
Routinierter Umgang mit Git / CI/CD und modernen Modellierungskonzepten (Star-Schema).
-
Mindset & Sprachen:
-
Pragmatische Hands-on-Mentalität, starke Kommunikationsfähigkeit und sehr gutes Deutsch & Englisch.
-
-
-
-
-
- Nachhaltig unterwegs mit unserem E-Auto-Leasing

