Rodeo
Get started

Sundayy

Data Engineer

England
Posted about 16 hours ago
Sign up to applySee more jobs like this
Get notified of more jobs like this · No spam, ever

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

About The Company

Medidata is at the forefront of transforming healthcare through innovative digital solutions that support clinical trials worldwide. With over 25 years of pioneering technological advancements, Medidata has successfully contributed to more than 38,000 trials involving 12 million patients. Our industry-leading platform combines advanced analytics, comprehensive data insights, and a vast clinical trial dataset to accelerate the development of life-saving therapies. Trusted by over 1 million registered users across approximately 2,300 customers, our seamless end-to-end solutions enhance patient experiences, streamline clinical processes, and expedite bringing new treatments to market. As a proud member of Dassault Systèmes (Euronext Paris: FR0014003TT8, DSY.PA), headquartered in New York City, Medidata has been recognized as a leader by Everest Group and IDC. Learn more at www.medidata.com and stay connected through our latest podcasts and social channels.

About The Role

We are seeking a highly skilled Principal Data Engineer / Architect to join our Data Platform Team in a hybrid remote/in-office capacity. Reporting directly to the Director of Engineering, you will lead the strategic development and hands-on implementation of our next-generation Object-Centric Data Fabric. Your primary focus will be to transition traditional application-centric architectures into a centralized semantic layer that unifies multi-stream operational data, including Electronic Data Capture (EDC), patient telemetry, and real-world health datasets. This role involves designing scalable, high-performance data lake architectures centered around Apache Iceberg, Snowflake, and heterogeneous compute engines, with a strong emphasis on AI augmentation to automate routine tasks such as schema inference, query optimization, and regulatory documentation. You will play a critical role in shaping the future of our clinical data infrastructure, enabling advanced AI/ML capabilities, and ensuring compliance with healthcare regulations. Your expertise will directly impact the speed and accuracy of clinical trial data processing, ultimately accelerating the development of new therapies and improving patient outcomes.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Data Science, Software Engineering, or related field.
  • Proven professional experience in enterprise data engineering and architecture, with a focus on data warehousing and data lake solutions.
  • Expert-level mastery of SQL (OLAP/OLTP) and cloud data platforms, particularly Snowflake.
  • Hands-on experience designing and implementing open data lake architectures using Apache Iceberg.
  • Strong proficiency in building production backend services in Java or Scala, preferably on AWS cloud infrastructure.
  • Deep familiarity with real-time streaming architectures such as Apache Kafka.
  • Practical experience integrating AI developer tools (e.g., GitHub Copilot, Cursor) into daily workflows to enhance coding efficiency.
  • Solid understanding of CI/CD pipelines, Git version control, TDD/BDD methodologies, and AI-driven testing frameworks.
  • Knowledge of clinical trial data workflows, healthcare data models, and compliance standards like GxP, HIPAA, and GDPR.
  • Ability to apply AI/ML tools for schema mapping, data validation, and ensuring zero-loss data integrity.

Responsibilities

  • Architect and evolve the enterprise semantic data fabric, transforming multi-stream clinical datasets into an object-centric model.
  • Design and implement a modern Data Lake strategy utilizing Apache Iceberg, ensuring high-performance querying and interoperability with Snowflake and other compute engines.
  • Develop and maintain automated, AI-augmented multi-stream ingestion pipelines for complex clinical trial schemas.
  • Build scalable, fault-tolerant real-time data ingestion pipelines using Kafka, AWS, and Snowflake.
  • Write robust backend services in Java or Scala, leveraging AI coding assistants for rapid development and performance tuning.
  • Promote best practices in database optimization, change data capture, data migration, and data aggregation techniques.
  • Integrate automated AI inferencing and reasoning layers into data pipelines to detect telemetry anomalies and map safety signals.
  • Lead initiatives in TDD and BDD, utilizing AI tools to generate synthetic datasets and validate compliance with regulatory standards.
  • Troubleshoot complex production issues across distributed data environments and design resilient, self-healing data pipelines.
  • Collaborate with cross-functional teams to align technical strategies with business goals, ensuring scalable and compliant data solutions.

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

Benefits

  • Comprehensive medical, dental, and vision insurance plans.
  • Life and disability insurance coverage.
  • Generous pension and retirement plans.
  • Over 25 paid holidays annually.
  • Opportunities for professional growth and development in a cutting-edge technological environment.
  • Flexible hybrid work arrangements to promote work-life balance.
  • Access to innovative tools and AI-driven workflows to enhance productivity.

Equal Opportunity

Medidata is committed to creating a diverse and inclusive workplace. We are an equal opportunity employer and do not discriminate based on race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, or veteran status. We encourage applicants of all backgrounds to apply and join our mission to improve healthcare through innovation.

Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Skills

SQL
Snowflake
Apache Iceberg
Java
Scala
AWS
Apache Kafka
Data Architecture
CI/CD
TDD/BDD
GxP
HIPAA
GDPR
AI-Augmented Development
Data Lake Design
Semantic Data Fabric

Location

England, United Kingdom

Sign up to applySee more jobs like this