09 ago
|
SoftServe
|
Chile
About The Role In this role, you will combine deep hands-on expertise in GPU-accelerated AI, Generative AI, and modern AI infrastructure with the opportunity to shape enterprise AI strategies for global clients. You will lead technical engagements from discovery through production, influence architectural decisions, drive AI adoption, and contribute to NVIDIA-focused go-to-market initiatives while collaborating with client stakeholders and multidisciplinary engineering teams.
Responsibilities Lead end-to-end AI engagements, from discovery and solution strategy through architecture design, implementation planning, and production delivery
Translate complex business challenges into AI use cases, solution roadmaps, and scalable enterprise architectures
Design and validate production-ready AI solutions leveraging NVIDIA technologies across cloud and on-premises environments
Define reference architectures for Generative AI and Agentic AI solutions, including GPU infrastructure, Kubernetes orchestration, and inference optimization strategies
Evaluate and optimize AI inference performance by analyzing GPU utilization, latency, throughput, scalability, and infrastructure efficiency
Benchmark and optimize LLM serving frameworks and deployment configurations
Lead pre-sales activities, including technical discovery, workshops, solution positioning, proposals, and proof-of-concept initiatives
Collaborate with NVIDIA stakeholders and internal teams to develop reusable accelerators, solution blueprints, and industry offerings
Drive technical thought leadership through whitepapers, technical content, conference presentations, and mentorship Requirements 6+ years of experience in AI consulting, Generative AI, Agentic AI, Machine Learning, or Deep Learning, including ownership of client-facing engagements
Bachelor's or Master's degree in Computer Science,
Applied Mathematics, Physics, Engineering, or related technical field preferred
Advanced expertise in Generative AI, Agentic AI, multimodal AI, transformers, Large Language Models (LLMs), and Vision Language Models (VLMs)
Hands-on experience with Python and modern AI/ML frameworks, including PyTorch, TensorFlow, Hugging Face, Pandas, and NumPy
Strong experience with NVIDIA AI technologies, including at least three of the following: NeMo, NIM, Triton, TensorRT-LLM, Riva, DeepStream, Metropolis, or Omniverse
Practical experience in deploying AI workloads on Kubernetes using Helm, NVIDIA GPU Operator, GPU device plugins, MIG/vGPU partitioning, and modern inference platforms such as vLLM or Ollama
Working knowledge of model quantization, inference optimization, and GPU profiling tools, including NVIDIA Nsight Systems, Nsight Compute, DCGM, PyTorch Profiler, and Triton or vLLM monitoring
Proven skill in analyzing GPU performance, identifying compute, memory, or I/O bottlenecks, and optimizing AI infrastructure for performance and cost efficiency
Experience in designing and deploying enterprise AI solutions on AWS, Azure, or GCP using CUDA, TensorRT, Triton Inference Server, DeepStream, and ONNX
Solid understanding of enterprise architecture, distributed systems, MLOps, AI governance, and modern software engineering practices
Strong advisory and stakeholder management skills
English proficiency for leading technical discussions with integral clients and stakeholders SoftServe is an equal opportunity employer. Qualified applicants will receive consideration regardless of race, color, ancestry, ethnicity, national origin, religion, sex, sexual orientation, gender identity or expression, age, citizenship, disability, health condition, marital or family status, veteran status, or any other characteristic protected by applicable law.
📌 Lead Data Scientist (Nvidia) (Chile)
🏢 SoftServe
📍 Chile