Sundayy
Data Engineer

How your CV stacks up
Upload your CV to see how well it fits this job role
?%
About Our Company
Medidata is a leading innovator in digital solutions that power smarter treatments and healthier lives through advanced clinical trial technology. With over 25 years of pioneering technological advancements, Medidata has contributed to more than 38,000 trials involving 12 million patients worldwide. Our industry-leading platform integrates analytics-powered insights and one of the largest clinical trial datasets in the industry, trusted by over 1 million registered users across approximately 2,300 customers. Headquartered in New York City and a proud brand of Dassault Systèmes, Medidata has been recognized as a leader by Everest Group and IDC for its innovative approach. Our mission is to accelerate the development of life-saving therapies by transforming raw data into high-impact insights, ultimately improving patient outcomes and streamlining clinical processes.
About The Role
We are seeking a highly skilled Principal Data Engineer / Architect to join our Data Platform Team in a hybrid remote/in-office capacity. Reporting directly to the Director of Engineering, you will be responsible for leading the strategic vision and hands-on development of our next-generation Object-Centric Data Fabric. Your primary focus will be on transforming traditional application-centric architectures into a centralized semantic layer that unifies multi-stream operational data, including Electronic Data Capture (EDC), patient telemetry, and real-world health datasets. This role involves designing and implementing a modern Data Lake architecture centered on Apache Iceberg, ensuring high-performance querying and seamless integration with Snowflake and various compute engines. You will leverage AI augmentation to automate routine tasks such as schema inference, query optimization, and regulatory documentation, allowing you to focus on high-impact platform architecture and innovation.
You will architect and evolve enterprise semantic data models, develop multi-stream ingestion pipelines, and build fault-tolerant real-time ingestion systems using Kafka, AWS, and Snowflake. Your expertise will guide technical strategy, promote best practices across teams, and optimize complex SQL workloads. Additionally, you will integrate automated AI inferencing and reasoning layers into data pipelines to detect anomalies and automatically map safety signals, enhancing real-time clinical trial monitoring. Leading initiatives around data quality, compliance, and system resilience will be key components of your role, ensuring the robustness and regulatory adherence of our data platform. Your leadership will directly contribute to advancing life sciences technology, enabling faster, safer, and more efficient clinical trials worldwide.
Reasons to use Rodeo
I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?
Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.
Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.
Start with a chat, not a search bar
Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.
Graduate Consultant — 2026 Scheme
Why you're a good match
StrongYour economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.
See breakdownIt searches the market for you
Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.
Why you're a good match
You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.
Experience fit
Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.
Only hits
No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.
Qualifications
- Bachelor’s or Master’s degree in Computer Science, Data Science, Software Engineering, or a related field.
- Proven professional experience in enterprise data engineering and architecture.
- Expertise in SQL (OLAP/OLTP) and cloud data platforms, especially Snowflake.
- Hands-on experience designing open data lake structures using Apache Iceberg.
- Proficiency in building production backend services using Java or Scala on AWS.
- Deep familiarity with real-time streaming architectures such as Apache Kafka.
- Experience integrating AI developer tools (e.g., GitHub Copilot, Cursor) into engineering workflows.
- Strong understanding of CI/CD pipelines, Git version control, and TDD/BDD practices.
- Knowledge of clinical trial data workflows, healthcare data models, and compliance standards (GxP, HIPAA, GDPR).
- Ability to leverage AI/ML tools for schema mapping, data validation, and performance optimization.
Responsibilities
- Lead the architecture and evolution of the enterprise semantic data fabric, transforming multi-stream clinical datasets into an object-centric model.
- Design and implement a modern Data Lake strategy centered on Apache Iceberg, ensuring high-performance querying and interoperability with Snowflake.
- Develop and maintain end-to-end ingestion pipelines for complex clinical trial schemas, utilizing AI-driven schema inference and ontology alignment tools.
- Build scalable, fault-tolerant real-time ingestion pipelines using Kafka, AWS, and Snowflake.
- Write robust enterprise backend services in Java or Scala, leveraging AI-assisted coding tools for rapid development and performance tuning.
- Promote best practices in database optimization, including CDC, clustering, data migration, and workload auto-tuning.
- Integrate automated AI inferencing and reasoning into data pipelines to detect telemetry anomalies and map safety signals in real-time.
- Lead initiatives in data quality, compliance, and validation using AI-generated test cases and documentation artifacts.
- Troubleshoot complex production issues, ensuring system resilience and implementing self-healing pipeline architectures.
- Collaborate with cross-functional teams to align technical strategies with organizational goals and regulatory requirements.


Get help with your application
Your very own career expert that helps elevate your application to the next level.
Benefits
- Comprehensive medical, dental, and vision insurance plans.
- Life and disability insurance coverage.
- Generous pension and retirement plans.
- Over 25 paid holidays annually.
- Opportunities for professional growth and development.
- Flexible hybrid work environment to support work-life balance.
- Access to cutting-edge technology and AI tools to enhance your engineering capabilities.
Equal Opportunity
Medidata is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, age, disability, or veteran status. We believe that diverse teams drive innovation and are essential to our mission of transforming life sciences through technology.
“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”
Jessica, London
Location