DevJobs

Senior C++ Engineer – Embedded Performance & AI Acceleration

Overview
Skills
  • C C ꞏ 6y
  • C++ C++ ꞏ 6y
  • Algorithm implementation
  • Vector processing
  • SIMD
  • Real-time systems
  • Performance profiling
  • Memory hierarchy
  • Firmware
  • Embedded systems
  • DSP
  • Caches
  • Algorithm optimization
  • CUDA
  • Computer vision
  • LLMs
  • Compiler development
  • Neural network inference optimization
  • OpenCL
  • CNNs
  • SDK development
  • Signal processing
  • Toolchain development
  • AI accelerator programming
  • Vision Transformers

Senior C++ Engineer – Embedded Performance & AI Acceleration

GSI Technology · Tel Aviv District, Israel (Hybrid)


Most AI roles ask you to call PyTorch APIs.

This one asks you to understand what's happening beneath them.


At GSI Technology (NASDAQ: GSIT), we're building the Gemini® APU — a compute-in-memory processor purpose-built to accelerate LLMs, vision models, and advanced signal processing with fundamentally different silicon.

This role fits two kinds of engineers: those coming from real-time / embedded systems who want their next challenge to be AI silicon instead of another firmware stack — and those with an algorithms or performance-optimization background who want to get closer to the metal. If either sounds like you, keep reading.

There's no existing playbook for what we're doing.

🔍 Why this role is different

The gap between modern AI models and novel hardware doesn't close itself.

You'll be the engineer who closes it — by working where C++, computer architecture, and real AI workloads intersect at a level most engineers never reach.

You won't be fine-tuning models. You'll be deciding how they execute on hardware that didn't exist two years ago. No prior AI/ML experience required — what matters is that you think natively in cache lines, memory access patterns, and hardware constraints.

⚙️ What you'll build

  • Implement and optimize AI workloads in modern C++ — from Python reference models to high-performance runtime implementations
  • Adapt LLMs, CNNs, and Vision Transformers to a unique compute-in-memory execution model
  • Identify and eliminate memory access, latency, and throughput bottlenecks
  • Design software libraries and runtime infrastructure for AI execution on novel silicon
  • Work directly with Hardware Architects — and shape future chip capabilities through software-driven insights

✅ What we need

  • Strong C/C++ — you're comfortable thinking in cache lines and memory access patterns
  • Experience close to hardware: embedded systems, DSP, real-time systems, or low-level firmware
  • Solid grasp of: memory hierarchy, caches, SIMD/vector processing, performance profiling
  • Background in algorithm implementation or optimization (signal processing, computer vision, or similar) — from any domain, not necessarily AI
  • 6+ years of software development experience
  • B.Sc. in Computer Science, Electrical Engineering, or equivalent hands-on experience
  • Sharp debugger. Fast learner. High ownership.

⭐ Strong bonus if you bring

  • Neural network inference optimization experience
  • CUDA, OpenCL, or AI accelerator programming
  • Compiler, SDK, or toolchain development
  • Familiarity with LLMs, Vision Transformers, or CNNs

📍 Tel Aviv, Ramat Hahayal | Full-Time

💰 Competitive compensation + (NASDAQ: GSIT)

  • Curious about the technical problem before you're ready to apply? Reach out — we're happy to start with a conversation.


Our Privacy Policy: Your resume and information will be kept confidential

GSI Technology