Location: Mississauga, ON Work Model: Hybrid Employment Type: Full-Time (FTE) Experience: 6 10 Years
Job Summary: We are seeking an experienced Kafka Administrator to manage, maintain, and optimize enterprise Apache Kafka environments. The ideal candidate will have strong expertise in Kafka administration, cluster management, performance tuning, security, monitoring, and production support within large-scale distributed systems.
Key Responsibilities
Install, configure, administer, and maintain Apache Kafka clusters across Development, QA, UAT, and Production environments.
Manage Kafka brokers, topics, partitions, replication factors, and consumer groups to ensure high availability and optimal performance.
Monitor Kafka cluster health, resource utilization, and capacity using enterprise monitoring tools.
Perform performance tuning, capacity planning, and cluster scaling to support growing workloads.
Troubleshoot Kafka infrastructure, messaging, connectivity, and performance-related issues, and perform root cause analysis.
Configure and manage Kafka security, including SSL/TLS, SASL authentication, ACLs, and encryption.
Administer Kafka Connect, Schema Registry, and related ecosystem components.
Implement and support High Availability (HA) and Disaster Recovery (DR) solutions.
Execute Kafka upgrades, patching, migrations, and platform maintenance with minimal business impact.
Collaborate with application, infrastructure, DevOps, and data engineering teams to onboard and support Kafka-based applications.
Develop and maintain operational runbooks, support documentation, and architecture documentation.
Configure monitoring, logging, alerting, and incident management processes.
Support production deployments, release activities, and environment readiness.
Automate operational and administrative tasks using Shell or Python scripting.
Ensure compliance with enterprise security, governance, audit, and regulatory standards.
Participate in production support and on-call rotation for critical incidents.
Drive continuous improvements to enhance platform reliability, scalability, security, and operational efficiency.