About the Company
Navitas Business Consulting, founded in 2006, has grown to become an industry leader in the digital transformation space. The company serves as a trusted advisor supporting a diverse client base across commercial, federal, and state & local markets. At their core, Navitas is a group of problem solvers providing award-winning technology solutions to drive digital acceleration for their customers. With proven solutions, award-winning technologies, and a team of expert problem solvers, Navitas consistently empowers customers to use technology as a competitive advantage and deliver cutting-edge transformative solutions.
Role Overview
We are seeking a highly skilled Databricks Engineer to design, build, and operate a modern Data & AI platform. This critical role focuses on establishing a strong foundation using the Medallion Architecture, which includes raw/bronze, curated/silver, and mart/gold layers. The successful candidate will be responsible for orchestrating complex data workflows and scalable ELT pipelines to integrate data from key enterprise systems such as PeopleSoft, D2L, and Salesforce. The ultimate goal is to deliver high-quality, governed data that powers machine learning, AI/BI, and analytics at scale. You will play an instrumental role in engineering the infrastructure and workflows that enable seamless data flow across the enterprise, ensuring operational excellence and providing the backbone for strategic decision-making, predictive modeling, and innovation.
Key Responsibilities
-
Data & AI Platform Engineering (Databricks-Centric):
- Design, implement, and optimize end-to-end data pipelines on Databricks, following Medallion Architecture principles.
- Build robust and scalable ETL/ELT pipelines using Apache Spark and Delta Lake to transform raw (bronze) data into trusted curated (silver) and analytics-ready (gold) data layers.
- Operationalize Databricks Workflows for orchestration, dependency management, and pipeline automation.
- Apply schema evolution and data versioning to support agile data development.
-
Platform Integration & Data Ingestion:
- Connect and ingest data from enterprise systems such as PeopleSoft, D2L, and Salesforce using APIs, JDBC, or other integration frameworks.
- Implement connectors and ingestion frameworks that accommodate structured, semi-structured, and unstructured data.
- Design standardized data ingestion processes with automated error handling, retries, and alerting.
-
Data Quality, Monitoring, and Governance:
- Develop data quality checks, validation rules, and anomaly detection mechanisms to ensure data integrity across all layers.
- Integrate monitoring and observability tools (e.g., Databricks metrics, Grafana) to track ETL performance, latency, and failures.
- Implement Unity Catalog or equivalent tools for centralized metadata management, data lineage, and governance policy enforcement.
-
Security, Privacy, and Compliance:
- Enforce data security best practices including row-level security, encryption at rest/in transit, and fine-grained access control via Unity Catalog.
- Design and implement data masking, tokenization, and anonymization for compliance with privacy regulations (e.g., GDPR, FERPA).
- Work with security teams to audit and certify compliance controls.
-
AI/ML-Ready Data Foundation:
- Enable data scientists by delivering high-quality, feature-rich data sets for model training and inference.
- Support AIOps/MLOps lifecycle workflows using MLflow for experiment tracking, model registry, and deployment within Databricks.
- Collaborate with AI/ML teams to create reusable feature stores and training pipelines.
-
Cloud Data Architecture and Storage:
- Architect and manage data lakes on Azure Data Lake Storage (ADLS) or Amazon S3, and design ingestion pipelines to feed the bronze layer.
- Build data marts and warehousing solutions using platforms like Databricks.
- Optimize data storage and access patterns for performance and cost-efficiency.
-
Documentation & Enablement:
- Maintain technical documentation, architecture diagrams, data dictionaries, and runbooks for all pipelines and components.
- Provide training and enablement sessions to internal stakeholders on the Databricks platform, Medallion Architecture, and data governance practices.
- Conduct code reviews and promote reusable patterns and frameworks across teams.
-
Reporting and Accountability:
- Submit a weekly schedule of hours worked and progress reports outlining completed tasks, upcoming plans, and blockers.
- Track deliverables against roadmap milestones and communicate risks or dependencies.
Tech Stack
- Data Platform: Databricks, Delta Lake, Apache Spark
- Programming Languages: SQL, Python, Scala
- Cloud: Azure (Azure Data Lake Storage - ADLS), AWS (Amazon S3)
- Integration: APIs, JDBC, Enterprise platforms (PeopleSoft, Salesforce, D2L)
- Data Governance & ML: Unity Catalog, MLflow
- Monitoring: Grafana, Databricks metrics
- Schema & Modeling: Star schema, Snowflake schema, Medallion Architecture (Bronze/Silver/Gold)