Optimize AMD GPU kernels from GGUF profiles to generate llama.cpp runtime configs with no recompilation and faster decode speed