make: Entering directory '/mnt/fileserver/prj/dmctp/user/src/tensor/src/dev_tools/profiler' make: Nothing to be done for 'all'. make: Leaving directory '/mnt/fileserver/prj/dmctp/user/src/tensor/src/dev_tools/profiler' Tensor profiler: 597 events, 16 graph nodes Summary capture span 590.168 ms first-to-last kernel span 317.322 ms kernel work (sum of kernel calls) 131.040 ms library work (inclusive, overlaps) 402.172 ms JIT compile work 239.455 ms Kernel execution operation dtype calls total_ms avg_ms MATRIX_INV_SQUARE REAL32 1 0.035 0.035 MATRIX_MUL REAL32 1 0.593 0.593 BLC UINT16 1 7.771 7.771 DPC UINT16 1 22.206 22.206 AWB UINT16 1 9.362 9.362 DEMOSAIC UINT16 1 19.094 19.094 CCM UINT16 1 15.858 15.858 LUT UINT16 1 12.431 12.431 BLUR UINT16 1 17.758 17.758 BLEND_RGB2YUV UINT16 1 18.799 18.799 CHROMA_SUBSAMPLE UINT16 1 7.132 7.132 Library tasks (inclusive time) task dtype calls total_ms avg_us bytes mem_pool_init UNKNOWN 1 0.004 3.657 48 mem_alloc UNKNOWN 21 0.182 8.673 565069856 mem_view_copy UNKNOWN 4 0.027 6.868 262252 mem_pool_plan UNKNOWN 16 0.096 6.008 538549684 mem_pool_set_cleanup UNKNOWN 16 0.044 2.765 0 tensor_lazy_alloc REAL32 6 0.215 35.815 262324 mem_free UNKNOWN 32 0.111 3.462 565069856 tensor_lazy_op_dispatch REAL32 2 55.696 27847.769 0 tensor_alloc REAL32 6 0.076 12.657 262324 mem_copy UNKNOWN 4 0.154 38.456 262252 tensor_view_to REAL32 1 0.003 2.825 36 tensor_lazy_view_to REAL32 1 0.779 779.476 36 mem_view UNKNOWN 2 0.019 9.703 52515840 tensor_lazy_alloc UINT16 10 0.246 24.568 538287360 tensor_lazy_op_dispatch UINT16 9 184.559 20506.602 0 tensor_alloc UINT16 10 0.181 18.087 538287360 tensor_view_to UINT16 2 0.006 3.221 118160640 tensor_lazy_view_to UINT16 2 131.354 65676.853 118160640 tensor_free UINT16 10 0.081 8.118 538287360 tensor_free REAL32 6 0.047 7.886 262324 tensor_lazy_shutdown UNKNOWN 1 0.217 216.513 0 mem_pool_shutdown UNKNOWN 1 28.075 28075.003 577937408 JIT cache hits: 13 cache misses: 11 hit rate: 54.2% task dtype calls total_ms avg_ms bitcode_bytes JIT PREPARE CACHE LOOKUP REAL32 2 0.014 0.007 0 tensor_op_00040002_8_8 REAL32 1 34.742 34.742 134884 tensor_op_00040001_8_8 REAL32 1 20.754 20.754 134884 JIT CACHE LOOKUP REAL32 2 0.013 0.007 0 JIT PREPARE CACHE LOOKUP UINT16 11 0.078 0.007 0 tensor_op_00010003_5_5 UINT16 1 17.632 17.632 134884 tensor_op_00010004_5_5 UINT16 1 25.370 25.370 134884 tensor_op_00010005_5_5 UINT16 1 17.250 17.250 134884 tensor_op_00010006_5_5 UINT16 1 23.301 23.301 134884 tensor_op_00010007_5_5 UINT16 1 19.558 19.558 134884 tensor_op_00010008_5_5 UINT16 1 16.157 16.157 134884 tensor_op_00010009_5_5 UINT16 1 24.484 24.484 134884 tensor_op_0001000a_5_5 UINT16 1 18.623 18.623 134884 tensor_op_0001000b_5_5 UINT16 1 21.584 21.584 134884 JIT CACHE LOOKUP UINT16 9 0.065 0.007 0 Memory traffic and lazy pool external heap allocations 5 26520172 bytes external heap storage releases 5 26520172 bytes memory-block handle releases 32 565069856 logical bytes copy operations 8 524504 bytes moved pool block allocations 16 pool planned bytes 538549824 bytes pool committed bytes 577937408 bytes pool bump high-water 538549824 bytes pool reserved but unused 39387584 bytes committed-space utilization 93.2% plan realization 100.0% Operation graph node operation op_code dtype runs kernel_ms parents N0 ALLOC 0x00000000 REAL32 1 0.000 [] N1 ALLOC 0x00000000 REAL32 1 0.000 [] N2 MATRIX_INV_SQUARE 0x00040002 REAL32 1 0.035 [N0] N3 MATRIX_MUL 0x00040001 REAL32 1 0.593 [N1,N2] N4 ALLOC 0x00000000 REAL32 1 0.000 [] N5 ALLOC 0x00000000 UINT16 1 0.000 [] N6 BLC 0x00010003 UINT16 1 7.771 [N5] N7 DPC 0x00010004 UINT16 1 22.206 [N6] N8 AWB 0x00010005 UINT16 1 9.362 [N7] N9 DEMOSAIC 0x00010006 UINT16 1 19.094 [N8] N10 CCM 0x00010007 UINT16 1 15.858 [N9,N4] N11 ALLOC 0x00000000 REAL32 1 0.000 [] N12 LUT 0x00010008 UINT16 1 12.431 [N10,N11] N13 BLUR 0x00010009 UINT16 1 17.758 [N12] N14 BLEND_RGB2YUV 0x0001000a UINT16 1 18.799 [N12,N13] N15 CHROMA_SUBSAMPLE 0x0001000b UINT16 1 7.132 [N14] kernel sequence (first execution of each node): N2:MATRIX_INV_SQUARE -> N3:MATRIX_MUL -> N6:BLC -> N7:DPC -> N8:AWB -> N9:DEMOSAIC -> N10:CCM -> N12:LUT -> N13:BLUR -> N14:BLEND_RGB2YUV -> N15:CHROMA_SUBSAMPLE Algorithmic throughput and tensor traffic kernel dtype class GOP/s tensor_GB/s OP/byte MATRIX_INV_SQUARE REAL32 floating-point 0.003 0.002 1.5000 MATRIX_MUL REAL32 floating-point 0.000 0.000 0.5000 BLC UINT16 integer 6.758 6.758 1.0000 DPC UINT16 integer 4.129 2.365 1.7461 AWB UINT16 integer 11.219 5.609 2.0000 DEMOSAIC UINT16 integer 11.001 5.501 2.0000 CCM UINT16 integer 17.386 9.935 1.7500 LUT UINT16 integer 28.516 12.695 2.2463 BLUR UINT16 integer 31.016 8.872 3.4961 BLEND_RGB2YUV UINT16 integer 26.538 12.571 2.1111 CHROMA_SUBSAMPLE UINT16 integer 4.602 16.568 0.2778 OP rates use the selected .op algorithm's static operations hint. tensor_GB/s is input+output tensor traffic, not hardware DRAM bandwidth. roofline data: tmp/isp_full_cpu_20260727.csv.roofline.csv roofline graph: tmp/isp_full_cpu_20260727.csv.roofline.svg kernel flame data: tmp/isp_full_cpu_20260727.csv.flame.folded