$190,000 – $280,000
Listed on Together AI’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
Together AI is hiring a Senior Network Engineer to design and operate global infrastructure for their AI compute platforms, working across multiple data centers and vendor networks. This role suits experienced network engineers comfortable troubleshooting complex, large-scale systems and collaborating across teams to solve problems that span networking, infrastructure, and applications.
Our summary, not Together AI’s wording. The full posting is on their site.
Skills this role names
- Amazon Web Services (AWS)
- Ansible
- Border Gateway Protocol (BGP)
- Cisco
- Git
- Juniper
- Kubernetes
- Linux
- Microsoft Azure
- Nmap
- NVIDIA Triton Inference Server
- OSPF
- Python
- VLAN
- Wireshark
Log in to see which of these are already on your profile.
What they ask for
Required
- 8+ years designing, building, and supporting large-scale production data center, cloud, service-provider, or HPC networks
- Deep understanding of TCP/IP and experience with BGP, OSPF, VXLAN, EVPN, ECMP, and QoS
- Experience designing and supporting multi-tenant network environments using VRFs, VLANs, overlays, and policy-based segmentation
- Hands-on experience deploying and troubleshooting networks from Arista, Cisco, Juniper, or NVIDIA
- Strong troubleshooting with Wireshark, tcpdump, MTR, curl, nmap, and Linux networking utilities
- Ability to diagnose connectivity, latency, packet loss, routing, and performance issues across network, host, and application layers
- Experience developing or maintaining network automation using Python, Ansible, or similar tools
- Experience with Git-based software development lifecycle including branching, code review, CI/CD, and deployment
- Working knowledge of Kubernetes networking including pods, services, CNIs, and connectivity troubleshooting
- Foundational knowledge of RDMA networking and RoCE or InfiniBand
- Experience with cloud networking in AWS, GCP, or Azure
- Strong Linux administration and troubleshooting skills
Nice to have
- Hands-on experience deploying or operating RoCE or InfiniBand fabrics
- Experience supporting GPU clusters, HPC environments, distributed storage, or high-bandwidth latency-sensitive workloads
- Understanding of AI training and inference traffic patterns and network demands
- Experience operating networks spanning thousands of devices across multiple data centers and geographic regions
- Familiarity with AI-assisted engineering tools and ability to validate and safely deploy AI-generated automation