• Projects 11
  • Rating 5.0
  • Rating 1 788

Budget: 15000 USD Deadline: 14 days

We at SDEV have extensive experience in optimizing computations on GPUs and configuring infrastructure for high-load tasks. We achieve this through profiling CUDA cores and fine-tuning Docker containers for efficient use of hardware accelerators. We are ready to ensure a balance of performance and quality for your models.

  • Projects 10
  • Rating 5.0
  • Rating 1 756

Budget: 50 USD Deadline: 1 day

Hello. For the implementation of the project, I will focus on developing low-level optimizations and adapting architectural solutions for different types of accelerators, such as GPU and TPU, using CUDA/ROCm and frameworks like TensorRT or OpenVINO. A scalable container infrastructure based on Docker and Kubernetes will be created for efficient deployment and management of computations, ensuring an optimal balance of performance and quality for large language models. I have already successfully implemented similar projects for optimizing transformers and have ready scripts and templates to accelerate setup and benchmarking. I suggest discussing all implementation details, final budget, and timelines in private messages.