Sundayy
Data Engineer

How your CV stacks up
Upload your CV to see how well it fits this job role
?%
About The Company
Medidata is at the forefront of transforming healthcare through innovative digital solutions that support clinical trials worldwide. With over 25 years of pioneering technological advancements, Medidata has successfully contributed to more than 38,000 trials involving 12 million patients. Our industry-leading platform combines advanced analytics, comprehensive data insights, and a vast clinical trial dataset to accelerate the development of life-saving therapies. Trusted by over 1 million registered users across approximately 2,300 customers, our seamless end-to-end solutions enhance patient experiences, streamline clinical processes, and expedite bringing new treatments to market. As a proud member of Dassault Systèmes (Euronext Paris: FR0014003TT8, DSY.PA), headquartered in New York City, Medidata has been recognized as a leader by Everest Group and IDC. Learn more at www.medidata.com and stay connected through our latest podcasts and social channels.
About The Role
We are seeking a highly skilled Principal Data Engineer / Architect to join our Data Platform Team in a hybrid remote/in-office capacity. Reporting directly to the Director of Engineering, you will lead the strategic development and hands-on implementation of our next-generation Object-Centric Data Fabric. Your primary focus will be to transition traditional application-centric architectures into a centralized semantic layer that unifies multi-stream operational data, including Electronic Data Capture (EDC), patient telemetry, and real-world health datasets. This role involves designing scalable, high-performance data lake architectures centered around Apache Iceberg, Snowflake, and heterogeneous compute engines, with a strong emphasis on AI augmentation to automate routine tasks such as schema inference, query optimization, and regulatory documentation. You will play a critical role in shaping the future of our clinical data infrastructure, enabling advanced AI/ML capabilities, and ensuring compliance with healthcare regulations. Your expertise will directly impact the speed and accuracy of clinical trial data processing, ultimately accelerating the development of new therapies and improving patient outcomes.
Reasons to use Rodeo
I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?
Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.
Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.
Start with a chat, not a search bar
Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.
Graduate Consultant — 2026 Scheme
Why you're a good match
StrongYour economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.
See breakdownIt searches the market for you
Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.
Why you're a good match
You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.
Experience fit
Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.
Only hits
No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.
Qualifications
- Bachelor’s or Master’s degree in Computer Science, Data Science, Software Engineering, or related field.
- Proven professional experience in enterprise data engineering and architecture, with a focus on data warehousing and data lake solutions.
- Expert-level mastery of SQL (OLAP/OLTP) and cloud data platforms, particularly Snowflake.
- Hands-on experience designing and implementing open data lake architectures using Apache Iceberg.
- Strong proficiency in building production backend services in Java or Scala, preferably on AWS cloud infrastructure.
- Deep familiarity with real-time streaming architectures such as Apache Kafka.
- Practical experience integrating AI developer tools (e.g., GitHub Copilot, Cursor) into daily workflows to enhance coding efficiency.
- Solid understanding of CI/CD pipelines, Git version control, TDD/BDD methodologies, and AI-driven testing frameworks.
- Knowledge of clinical trial data workflows, healthcare data models, and compliance standards like GxP, HIPAA, and GDPR.
- Ability to apply AI/ML tools for schema mapping, data validation, and ensuring zero-loss data integrity.
Responsibilities
- Architect and evolve the enterprise semantic data fabric, transforming multi-stream clinical datasets into an object-centric model.
- Design and implement a modern Data Lake strategy utilizing Apache Iceberg, ensuring high-performance querying and interoperability with Snowflake and other compute engines.
- Develop and maintain automated, AI-augmented multi-stream ingestion pipelines for complex clinical trial schemas.
- Build scalable, fault-tolerant real-time data ingestion pipelines using Kafka, AWS, and Snowflake.
- Write robust backend services in Java or Scala, leveraging AI coding assistants for rapid development and performance tuning.
- Promote best practices in database optimization, change data capture, data migration, and data aggregation techniques.
- Integrate automated AI inferencing and reasoning layers into data pipelines to detect telemetry anomalies and map safety signals.
- Lead initiatives in TDD and BDD, utilizing AI tools to generate synthetic datasets and validate compliance with regulatory standards.
- Troubleshoot complex production issues across distributed data environments and design resilient, self-healing data pipelines.
- Collaborate with cross-functional teams to align technical strategies with business goals, ensuring scalable and compliant data solutions.


Get help with your application
Your very own career expert that helps elevate your application to the next level.
Benefits
- Comprehensive medical, dental, and vision insurance plans.
- Life and disability insurance coverage.
- Generous pension and retirement plans.
- Over 25 paid holidays annually.
- Opportunities for professional growth and development in a cutting-edge technological environment.
- Flexible hybrid work arrangements to promote work-life balance.
- Access to innovative tools and AI-driven workflows to enhance productivity.
Equal Opportunity
Medidata is committed to creating a diverse and inclusive workplace. We are an equal opportunity employer and do not discriminate based on race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, or veteran status. We encourage applicants of all backgrounds to apply and join our mission to improve healthcare through innovation.
“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”
Jessica, London
Skills
Location