feat: migrate Python bindings from pybind11 to nanobind - #243
Open
hush-coder wants to merge 698 commits into
Open
feat: migrate Python bindings from pybind11 to nanobind#243hush-coder wants to merge 698 commits into
hush-coder wants to merge 698 commits into
Conversation
hush-coder
commented
Sep 1, 2025
- Bindings: migrate pymlir, pyfinalmlir, pyruntime_* to nanobind (NB_MODULE, nb::ndarray<>, .def_ro/.def_rw).
- Build: switch to find_package(nanobind) + nanobind_add_module(... NB_STATIC); add find_package(Python ... Development) at top-level.
- Dependencies: update requirements.txt/Docker to install nanobind>=2.8.0.
- Compatibility: keep python/tpu_mlir/RegisterEverything.cpp on pybind11.
1. use Lut to calculate EXP(quant(x)) in softmax 2. add backend-api for local reduceOp. Change-Id: If7b6bef0ac6c1a614c3c6ae09e5643ea16760cc9
Change-Id: I512e398d7153fd2568646216a1bd4785d1137d06
--the existence of split pattern leads to the failure of qtable, this can fix above bug Change-Id: Iec98a0caec0d477cad56bdb0b7f18de6fc68ac7f
to avoid some warning Change-Id: I6775382b7a9848989e91258b6a1d8ad595640eae
assign more buffer_size in convbwd for bm1684x train Change-Id: I734bdc88c1d6f3c5764bd1e4c6e0c80e9ea73979
backend of convOp doesnot support F8 conv in multi-branch cases. Change-Id: I0e1b759ce2d98afcfa6376369b8769ececcdb69c
This reverts commit 89ddeed. Change-Id: I9ad462837249875f5dd33d0ce4b3b8a98d04d2c6
support_multi_core consider l2 buffer size as well Change-Id: I0fd85d86f677c4f8e7603d24f9cefe2994c4e487
w4a16 support sym/asym per group quant Change-Id: I2a4e1094976fd92b0d7d30f9d39efe2a79460cd8
1) support compile llm, llama or qwen2 2) fast by access llm weights directly Change-Id: Idfb7c7a2bd293dcd3dd89657de7e352de5116e79
1. v3.3 -> v3.4 2. align python lib version Change-Id: Ia7e2dded8f54cf9cdf7a0fd43b09e61da46b5ca4
1. This reverts commit d8a657b. 2. Only applicable to docker images starting from v3.4 Change-Id: I2fac2a6815e1eb973de47b4a35ab94a17a2f971a
compile bug caused by l2mem buffer calculate at lowering stage Change-Id: I334de940488cf2de5cbe240b8cce113f6b79f8ba
288982 -> 5feacd Change-Id: Ice334f647ad63c019f3234181ff4857aaa6d2146
fix mars3 regression Change-Id: Ie2ceea1b6cc6824482460e29ca5fe0ec4927b979
Change-Id: I64ca79e21db483fb61e5ca0321e0e714c49d2f8e
-- int4 and int8 mix quantization freely Change-Id: I9ee9be169dbbc3f4eb4470ed20a91b42bd96c9e6
1) llm_convert.py -m /workspace/Qwen2-7B-Instruct -s 384 -q w4bf16 -g 128 -c bm1684x --out_dir qwen2_7b 2) add doc to introduce this tool Change-Id: Id89bde486bd161ff7da5a2131490e172f919064a
test by: llm_convert.py -m /workspace/Qwen2.5-3B-Instruct -s 384 -q w4bf16 -g 128 -c bm1684x --out_dir qwen2.5_3b Change-Id: I928115e1510fe0a363d439a46642756c81d9cee2
add LutOp, BinaryConstShiftOp, BinaryShiftOp to group mm Change-Id: I929d2df66a9a2fefb94dde3581b3b6b18df52e74
Calculate the ranges of activations and weights. Then scale the weights of layernorm and matmuls after the layernorm. Change-Id: I1701247cc7f65da7b19548317c69cc47e17a6fcc
1.using multi maxpool to identify yolo models 2.postprocess block as F16 Change-Id: I69af2c7a804c1be8a3009766e422444067decca7
update backend Change-Id: Ie0bcfeb4a6995b7af2d52c077a7c7b920c3bdd11
- Adding the custom operator 'Correlation'. - The backend code uses PPL. - Adding ONNX Runtime call logic when custom operators are present. Change-Id: I9da52e51d8baa871e3cf5c603cf56b32a6e2a983
1. the header parse error 2. fix cmd group bin same file error Change-Id: Ice6c326bec08f688c86c7b6174d09fdc38306318
Change-Id: I84e75a3ded7f66d26f3c27118057f5fd168670a6
1.replace op name in regression calitable for c755a68 Change-Id: Ic75f04a29274a9dd2d9d6e9720dbeaa7cd32fb9e
tpulang add a16matmul interface && support a16matmul in top layer Change-Id: Ib9dc4f512f832341f8cd7a906b1bde4256665071
analysys all passes and auto generated a tpuc cmd builder (to replace preparing command manully in mlir_shell before) Change-Id: Ifc9e918bc048de38745c47bac5a295ca91c49cb6
fix w8a16 with group_size fail error Change-Id: Id3fa7121ff96e436d8315be7ea4594faafb1bbff
only pdate exist tensor_infos Change-Id: I1b2a0e120d20bc89f9eb2a6966b8e58290279e14
- doc modification will trigger sensitive words checking Change-Id: I0ad6cf11ba55ef15dc19bf5cb9a9a84de22ab52e
1.add lock for getOpParam function Change-Id: I630631f417a9170349e2cd31b12c7572b526b0b8
- layer norm axis=5 support group 3d Change-Id: I3d82dfcbbccfc6260d1c8cf02f9b188f876fefd3
…tension import request - mainly for hik release env, hik env doesn't have onnx related python package Change-Id: If37ff7233b8b87d2d07e3b74db833b26f70b2c13
- moveUnaryPermute for BinaryShiftOp, the second input should be weight support two same-order PermuteOp as one BinaryShiftOp's input later Change-Id: Ia4c49bd0efdcacbd1924e1868cadb3067084205b
-when concat input is inplaceOp,addr set error Change-Id: Ie361d0dae1d26736715a086a39e5f20345db346b
- fix AddOp shape infer bug if input is scalar - fix TopGatherToSlice pattern - shut down TopGatherToSliceByUnsqueeze pattern - add SubTryLowering - fix ArithOps' try lowering function Change-Id: I21074764b9af59d64c57e70916c516bee773c6d8
- fix Fattension highPrecision dyn ctrl && codegen issue - refile build.sh. Check the library's MD5 value instead of the 1686 commit ID Change-Id: If6bb9c7c334c53b7fd0a07624f93e8c0d1a4700c
- support ConcatOp shape inference when out_shape size > 1 Change-Id: If7ad8738bcab8dee2a19f36b388687038c726f89
- rename bmrt_test.py to prevent conflict with the bmrt_test tool in tpu-runtime Change-Id: I7187fd52ed99417e393d59331dce49e9fefcd402
- add llm_analyse.py info - add llm dtype check Change-Id: I95772ea8cabce1fb2bc81792e627500dd8faadc2
- script and model test case swap - disable test_fx Change-Id: Ifa8883dc3a72f59977a33431ebd7cd1ad04a6c00
-custom op for vmamba Change-Id: I88506ef0d54a2895ebf7ba7dc3645501c6e0d1da
-MLIR-581 -SimplifyRedundantCast ignore i32/8 to si32/8 Change-Id: I8f49902be4218f5ca7159ea6387dcd4b324d045d
This commit fixes the ConvertMatMul2Attention pattern by adding a check for dynamic mode to prevent failures when the module is in dynamic mode. Change-Id: Icab1891d2c5a46b4d5bc50c479257a30abb48c1d
- sync code with tpu-bmodel: add HAU and SDMA memory size for BM1690 middle buffer size only for BM1684 print core num Change-Id: Ic1698ff4db738716c494f6e87e632cb049e583de
- And, Xor, Or, Not Change-Id: Ieae37cfac4f5f0957e46dfd4b4b8e430737b1bc1
-FxGraphConverter support Indexput by scatternd -scatternd: support inplace_add fix buffer_size Change-Id: I6aa02a8913e8bcdffe71b1f5b9e52c94afb31e7f
- multi-core mm backend not support for bm1688 Change-Id: Idf72b00dfdbc0497dce2e406cfd1652ad092a1fb
-Reshape when the C_dims is too large Change-Id: I6dd4ccf395da6ea42ede01dfa894feb7ea2b0b08
- The TPUC_OPT_TIMEOUT environment variable can be used to set the timeout (unit:h/m/s). - Currently only applies to model_zoo regression. - Does not support _os_system_log. Change-Id: I62a52277cfaa09b56b517598e4a011582524aa14
- use options_.debugger to judge whether use manual group cost Change-Id: I56c41c70beded59ce05b6e9a872ff483665fc3b7
- share a long prompt for different questions Change-Id: I3a458970875b7d59ae34ea7902947f91e9b38545
-suitable for bmcv uint8 postprocess Change-Id: I8a7f83725d416eb3da315864794ee2bca487bb6e
检查nanobind
Author
|
抱歉,前段时间被其他工作绊住了 我在build.sh中加入了相关操作 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
