Docs & Rules

beehive-lab/TornadoVMGitHubLast refreshed Oct 3, 2026

This is a public, read-only report. Striff reads this repository's docs, turns each sentence that makes a claim about the code into a rule, and checks the rule against the code on the default branch. How this works

mastera86db1alisted 2h ago

Is this yours? Install to manage it

Once installed, Striff checks every pull request.

139
62
46
6
19
21
16
10
10
1
1
2
2
8
8
8
56
36 docs, listed Oct 3
beehive-lab/TornadoVMOpen repository

beehive-lab/TornadoVM — documented rules

139 of 139 rules, printed October 3, 2026.

The sentence in your docs
Read Oct 3
Concern Kernel-entry resolution, Where TaskUtils.resolveViaSerializedLambda, Notes lambda → SerializedLambda → Method
TaskUtils has a resolveViaSerializedLambda method
Holds
Read Oct 3
Concern SPI (no HotSpot type), Where TornadoBackendProvider.createBackend(OptionValues, TornadoVMConfigAccess), Notes backends construct reflection providers directly
TornadoBackendProvider has a createBackend method
Holds
Read Oct 3
OCLHotSpotBackendFactory.createJITCompiler → new TornadoMetaAccessProvider(), new TornadoConstantReflectionProvider(snippetReflection) (no HotSpotJVMCIRuntime param).
OCLHotSpotBackendFactory depends on TornadoMetaAccessProvider
Holds
Read Oct 3
OCLHotSpotBackendFactory.createJITCompiler → new TornadoMetaAccessProvider(), new TornadoConstantReflectionProvider(snippetReflection) (no HotSpotJVMCIRuntime param).
OCLHotSpotBackendFactory depends on TornadoConstantReflectionProvider
Holds
Read Oct 3
OCLHotSpotBackendFactory.createJITCompiler → new TornadoMetaAccessProvider(), new TornadoConstantReflectionProvider(snippetReflection) (no HotSpotJVMCIRuntime param).
OCLHotSpotBackendFactory has createJITCompiler
Holds
Read Oct 3
The *GraphBuilderPlugins.registerInvocationPlugins signature must take jdk.vm.ci.meta.MetaAccessProvider, not HotSpotMetaAccessProvider.
OCLGraphBuilderPlugins has registerInvocationPlugins
Holds
Read Oct 3
OCLFieldBuffer resolves the object type via TornadoCoreRuntime.getTornadoRuntime().getMetaAccess() and the reflection ResolvedJavaType (no HotSpotResolvedJavaType cast).
OCLFieldBuffer depends on TornadoCoreRuntime
Holds
Read Oct 3
graal/phases/TornadoOpenCLIntrinsicsReplacements (registered in OCLHighTier right after HighTierLoweringPhase).
OCLHighTier depends on TornadoOpenCLIntrinsicsReplacements
Holds
Read Oct 3
Also OCLLoweringProvider#isLocalArrayInput matches surviving contains("LocalArray") invokes at lowering time.
OCLLoweringProvider has isLocalArrayInput
Holds
Read Oct 3
OCLBackend.emitMethodParameters continues past a KernelContext/AtomicInteger parameter so it is not emitted as a __global buffer arg.
OCLBackend depends on KernelContext
Holds
Read Oct 3
OCLBackend.emitMethodParameters continues past a KernelContext/AtomicInteger parameter so it is not emitted as a __global buffer arg.
OCLBackend has emitMethodParameters
Holds
Read Oct 3
OCLNodeLIRBuilder.emitPrologue — a KernelContext/AtomicInteger parameter whose only remaining usages are FrameStates (deopt) must not load a buffer;
OCLNodeLIRBuilder has emitPrologue
Holds
Read Oct 3
OCLLoweringProvider uses (int) TornadoOptions.PANAMA_OBJECT_HEADER_SIZE for panama array element addresses;
OCLLoweringProvider depends on TornadoOptions
Holds
Read Oct 3
OCLAssembler.formatConstant escapes a TornadoObjectConstant wrapping a String as a C string literal (a HotSpotObjectConstant no longer appears on the reflection path).
OCLAssembler depends on TornadoObjectConstant
Holds
Read Oct 3
OCLAssembler.formatConstant escapes a TornadoObjectConstant wrapping a String as a C string literal (a HotSpotObjectConstant no longer appears on the reflection path).
OCLAssembler has formatConstant
Holds
Read Oct 3
OCLGraphBuilderPlugins.registerNativeArrayAccessPlugins → registerNativeArrayGetSet(Int/Float/ Double/Long/Short/Byte/Int8/Char) + registerHalfFloatArrayGetSet + the arch-neutral arrayElementAddress helper (&segment[(baseIndex+index)*elementBytes]).
OCLGraphBuilderPlugins has registerNativeArrayAccessPlugins
Holds
Read Oct 3
`cuDNN also exposes cudnnSoftmax, cudnnRelu/cudnnSigmoid/cudnnTanh, cudnnMaxPool2d, and fused FP16 scaled-dot-product (flash) attention via sdpaForward` — see the full provider catalog linked below.
CuDnn has cudnnSoftmax
Holds
Read Oct 3
`cuDNN also exposes cudnnSoftmax, cudnnRelu/cudnnSigmoid/cudnnTanh, cudnnMaxPool2d, and fused FP16 scaled-dot-product (flash) attention via sdpaForward` — see the full provider catalog linked below.
CuDnn has cudnnRelu
Holds
Read Oct 3
`cuDNN also exposes cudnnSoftmax, cudnnRelu/cudnnSigmoid/cudnnTanh, cudnnMaxPool2d, and fused FP16 scaled-dot-product (flash) attention via sdpaForward` — see the full provider catalog linked below.
CuDnn has cudnnSigmoid
Holds
Read Oct 3
`cuDNN also exposes cudnnSoftmax, cudnnRelu/cudnnSigmoid/cudnnTanh, cudnnMaxPool2d, and fused FP16 scaled-dot-product (flash) attention via sdpaForward` — see the full provider catalog linked below.
CuDnn has cudnnTanh
Holds
Read Oct 3
`cuDNN also exposes cudnnSoftmax, cudnnRelu/cudnnSigmoid/cudnnTanh, cudnnMaxPool2d, and fused FP16 scaled-dot-product (flash) attention via sdpaForward` — see the full provider catalog linked below.
CuDnn has cudnnMaxPool2d
Holds
Read Oct 3
`cuDNN also exposes cudnnSoftmax, cudnnRelu/cudnnSigmoid/cudnnTanh, cudnnMaxPool2d, and fused FP16 scaled-dot-product (flash) attention via sdpaForward` — see the full provider catalog linked below.
CuDnn has sdpaForward
Holds
Read Oct 3
Note that the TornadoVM profiler works only if enabled in the execution plan (via the `withProfiler` method).
TornadoExecutionPlan has a withProfiler method
Holds
Read Oct 3
**NOTE:** The methods `init() and clear()` are essential because, contrary to their counterpart primitive arrays which are initialized by default with 0, the new types contain garbage values when first created.
ByteArray has an init method
Holds
Read Oct 3
**NOTE:** The methods `init() and clear()` are essential because, contrary to their counterpart primitive arrays which are initialized by default with 0, the new types contain garbage values when first created.
ByteArray has a clear method
Holds
Read Oct 3
**NOTE:** The methods `init() and clear()` are essential because, contrary to their counterpart primitive arrays which are initialized by default with 0, the new types contain garbage values when first created.
CharArray has an init method
Holds
Read Oct 3
**NOTE:** The methods `init() and clear()` are essential because, contrary to their counterpart primitive arrays which are initialized by default with 0, the new types contain garbage values when first created.
CharArray has a clear method
Holds
Read Oct 3
**NOTE:** The methods `init() and clear()` are essential because, contrary to their counterpart primitive arrays which are initialized by default with 0, the new types contain garbage values when first created.
DoubleArray has an init method
Holds
Read Oct 3
**NOTE:** The methods `init() and clear()` are essential because, contrary to their counterpart primitive arrays which are initialized by default with 0, the new types contain garbage values when first created.
DoubleArray has a clear method
Holds
Read Oct 3
**NOTE:** The methods `init() and clear()` are essential because, contrary to their counterpart primitive arrays which are initialized by default with 0, the new types contain garbage values when first created.
FloatArray has an init method
HoldsPR #1069 · Sep 30
Read Oct 3
**NOTE:** The methods `init() and clear()` are essential because, contrary to their counterpart primitive arrays which are initialized by default with 0, the new types contain garbage values when first created.
FloatArray has a clear method
HoldsPR #1069 · Sep 30
Read Oct 3
**NOTE:** The methods `init() and clear()` are essential because, contrary to their counterpart primitive arrays which are initialized by default with 0, the new types contain garbage values when first created.
IntArray has an init method
Holds
Read Oct 3
**NOTE:** The methods `init() and clear()` are essential because, contrary to their counterpart primitive arrays which are initialized by default with 0, the new types contain garbage values when first created.
IntArray has a clear method
Holds
Read Oct 3
**NOTE:** The methods `init() and clear()` are essential because, contrary to their counterpart primitive arrays which are initialized by default with 0, the new types contain garbage values when first created.
LongArray has an init method
Holds
Read Oct 3
**NOTE:** The methods `init() and clear()` are essential because, contrary to their counterpart primitive arrays which are initialized by default with 0, the new types contain garbage values when first created.
LongArray has a clear method
Holds
Read Oct 3
**NOTE:** The methods `init() and clear()` are essential because, contrary to their counterpart primitive arrays which are initialized by default with 0, the new types contain garbage values when first created.
ShortArray has an init method
Holds
Read Oct 3
**NOTE:** The methods `init() and clear()` are essential because, contrary to their counterpart primitive arrays which are initialized by default with 0, the new types contain garbage values when first created.
ShortArray has a clear method
Holds
Read Oct 3
`local only means the operation is supported when the array was allocated with KernelContext.allocateIntLocalArray(int)`;
KernelContext has an allocateIntLocalArray method
Holds
Read Oct 3
The Task-Graph API also defines a method, named `transferToDevice` to set which arrays need to be copied to the target accelerator.
TaskGraph has a transferToDevice method
HoldsPR #1069 · Sep 30
Read Oct 3
Similar to `transferToDevice, the TaskGraph API also offers a call to sync the data back from the device to the host. The API call is transferToHost` with the following parameters:
TaskGraph has transferToHost
HoldsPR #1069 · Sep 30
Read Oct 3
TornadoVM supports batch processing through the following API call `withBatch of the TornadoExecutionPlan` API.
TornadoExecutionPlan has withBatch
HoldsPR #1069 · Sep 30
Read Oct 3
`Tile exposes getDType, getRank, getDimension and getElement` (JVM path only, for tests and host debugging);
Tile has getDType
HoldsPR #1069 · Sep 30
Read Oct 3
`Tile exposes getDType, getRank, getDimension and getElement` (JVM path only, for tests and host debugging);
Tile has getRank
HoldsPR #1069 · Sep 30
Read Oct 3
`Tile exposes getDType, getRank, getDimension and getElement` (JVM path only, for tests and host debugging);
Tile has getDimension
HoldsPR #1069 · Sep 30
Read Oct 3
`Tile exposes getDType, getRank, getDimension and getElement` (JVM path only, for tests and host debugging);
Tile has getElement
HoldsPR #1069 · Sep 30
Read Oct 3
`TensorView exposes getDType, getRank, getExtent, and getElementOffset/getRowStride` for a strided view;
TensorView has getDType
HoldsPR #1069 · Sep 30
Read Oct 3
`TensorView exposes getDType, getRank, getExtent, and getElementOffset/getRowStride` for a strided view;
TensorView has getRank
HoldsPR #1069 · Sep 30
Read Oct 3
`TensorView exposes getDType, getRank, getExtent, and getElementOffset/getRowStride` for a strided view;
TensorView has getExtent
HoldsPR #1069 · Sep 30
Read Oct 3
`TensorView exposes getDType, getRank, getExtent, and getElementOffset/getRowStride` for a strided view;
TensorView has getElementOffset
HoldsPR #1069 · Sep 30
Read Oct 3
`TensorView exposes getDType, getRank, getExtent, and getElementOffset/getRowStride` for a strided view;
TensorView has getRowStride
HoldsPR #1069 · Sep 30
Read Oct 3
`PartitionView adds getTileDimension and getBlockCount`.
PartitionView has getTileDimension
HoldsPR #1069 · Sep 30
Read Oct 3
`PartitionView adds getTileDimension and getBlockCount`.
PartitionView has getBlockCount
HoldsPR #1069 · Sep 30
Read Oct 3
`DType names F16, BF16, F32, F64, TF32, FP8_E4M3, FP8_E5M2, S8, S32 and the comparison-only PRED`.
DType has F16
HoldsPR #1069 · Sep 30
Read Oct 3
`DType names F16, BF16, F32, F64, TF32, FP8_E4M3, FP8_E5M2, S8, S32 and the comparison-only PRED`.
DType has BF16
HoldsPR #1069 · Sep 30
Read Oct 3
`DType names F16, BF16, F32, F64, TF32, FP8_E4M3, FP8_E5M2, S8, S32 and the comparison-only PRED`.
DType has F32
HoldsPR #1069 · Sep 30
Read Oct 3
`DType names F16, BF16, F32, F64, TF32, FP8_E4M3, FP8_E5M2, S8, S32 and the comparison-only PRED`.
DType has F64
HoldsPR #1069 · Sep 30
Read Oct 3
`DType names F16, BF16, F32, F64, TF32, FP8_E4M3, FP8_E5M2, S8, S32 and the comparison-only PRED`.
DType has TF32
HoldsPR #1069 · Sep 30
Read Oct 3
`DType names F16, BF16, F32, F64, TF32, FP8_E4M3, FP8_E5M2, S8, S32 and the comparison-only PRED`.
DType has FP8_E4M3
HoldsPR #1069 · Sep 30
Read Oct 3
`DType names F16, BF16, F32, F64, TF32, FP8_E4M3, FP8_E5M2, S8, S32 and the comparison-only PRED`.
DType has FP8_E5M2
HoldsPR #1069 · Sep 30
Read Oct 3
`DType names F16, BF16, F32, F64, TF32, FP8_E4M3, FP8_E5M2, S8, S32 and the comparison-only PRED`.
DType has S8
HoldsPR #1069 · Sep 30
Read Oct 3
`DType names F16, BF16, F32, F64, TF32, FP8_E4M3, FP8_E5M2, S8, S32 and the comparison-only PRED`.
DType has S32
HoldsPR #1069 · Sep 30
Read Oct 3
`DType names F16, BF16, F32, F64, TF32, FP8_E4M3, FP8_E5M2, S8, S32 and the comparison-only PRED`.
DType has PRED
HoldsPR #1069 · Sep 30
Read Oct 3
Factories (uk.ac.manchester.tornado.cublas.CuBlas):
CuBlas is in uk.ac.manchester.tornado.cublas
Holds
Read Oct 3
Factory cublasSgemv(op, m, n, alpha, A, lda, x, incx, beta, y, incy), Operation y = alpha·op(A)·x + beta·y
CuBlas has a cublasSgemv method
Holds
Read Oct 3
Factory cublasSgemm(opA, opB, m, n, k, alpha, A, lda, B, ldb, beta, C, ldc), Operation C = alpha·op(A)·op(B) + beta·C
CuBlas has a cublasSgemm method
Holds
Read Oct 3
Factory cublasSgemmTF32(...), Operation SGEMM using TF32 tensor cores
CuBlas has a cublasSgemmTF32 method
Holds
Read Oct 3
Factory cublasGemmExFP16(...), Operation FP16 inputs and output, tensor-core GEMM
CuBlas has a cublasGemmExFP16 method
Holds
Read Oct 3
Factory cublasGemmExFP16FP32(...), Operation FP16 inputs, FP32 output — keeps accumulation precision at the boundary
CuBlas has a cublasGemmExFP16FP32 method
Holds
Read Oct 3
Factory cublasGemmExBF16(...), Operation BF16 inputs and output ( BFloat16Array ), tensor-core GEMM
CuBlas has a cublasGemmExBF16 method
Holds
Read Oct 3
Factory cublasSgemmStridedBatched(...), Operation batched SGEMM
CuBlas has a cublasSgemmStridedBatched method
Holds
Read Oct 3
Factory ltMatmulFP32 / ltMatmulFP16, Epilogue none
CuBlasLt has ltMatmulFP32
Holds
Read Oct 3
Factory ltMatmulFP32 / ltMatmulFP16, Epilogue none
CuBlasLt has ltMatmulFP16
Holds
Read Oct 3
Factory ltMatmulFP8, Epilogue none (E4M3 operands, FP16 output, TN form, ld multiple of 16 B)
CuBlasLt has ltMatmulFP8
Holds
Read Oct 3
Factory ltMatmulBiasFP16, Epilogue BIAS
CuBlasLt has ltMatmulBiasFP16
Holds
Read Oct 3
Factory ltMatmulGeluBiasFP16, Epilogue GELU_BIAS (tanh approximation)
CuBlasLt has ltMatmulGeluBiasFP16
Holds
Read Oct 3
Factory cufftForwardC2C / cufftInverseC2C(input, output, n, batch), Transform 1D complex→complex
CuFft has cufftForwardC2C
Holds
Read Oct 3
Factory cufftForwardC2C / cufftInverseC2C(input, output, n, batch), Transform 1D complex→complex
CuFft has a cufftInverseC2C method
Holds
Read Oct 3
Factory cufftForwardR2C / cufftInverseC2R(input, output, n, batch), Transform real↔complex (Hermitian, n ↔ n/2+1 )
CuFft has cufftForwardR2C
Holds
Read Oct 3
Factory cufftForwardR2C / cufftInverseC2R(input, output, n, batch), Transform real↔complex (Hermitian, n ↔ n/2+1 )
CuFft has a cufftInverseC2R method
Holds
Read Oct 3
Factory cufftForwardZ2Z / cufftInverseZ2Z(input, output, n, batch), Transform FP64 complex→complex, 1D
CuFft has cufftForwardZ2Z
Holds
Read Oct 3
Factory cufftForwardZ2Z / cufftInverseZ2Z(input, output, n, batch), Transform FP64 complex→complex, 1D
CuFft has a cufftInverseZ2Z method
Holds
Read Oct 3
Factory cufftForward2dC2C / cufftInverse2dC2C(input, output, nx, ny), Transform 2D complex→complex
CuFft has cufftForward2dC2C
Holds
Read Oct 3
Factory cufftForward2dC2C / cufftInverse2dC2C(input, output, nx, ny), Transform 2D complex→complex
CuFft has a cufftInverse2dC2C method
Holds
Read Oct 3
Factory cudnnSoftmax(input, output, rows, cols), Operation per-row numerically-stable softmax
CuDnn has a cudnnSoftmax method
Holds
Read Oct 3
Factory cudnnRelu / cudnnSigmoid / cudnnTanh(input, output, size), Operation activations
CuDnn has a cudnnRelu method
Holds
Read Oct 3
Factory cudnnRelu / cudnnSigmoid / cudnnTanh(input, output, size), Operation activations
CuDnn has a cudnnSigmoid method
Holds
Read Oct 3
Factory cudnnRelu / cudnnSigmoid / cudnnTanh(input, output, size), Operation activations
CuDnn has a cudnnTanh method
Holds
Read Oct 3
Factory cudnnMaxPool2d(input, output, n, c, h, w, window, stride), Operation 2D max pooling
CuDnn has a cudnnMaxPool2d method
Holds
Read Oct 3
Factory cudnnConv2d(input, filter, output, n, c, h, w, k, r, s, pad, stride), Operation 2D convolution
CuDnn has a cudnnConv2d method
Holds
Read Oct 3
Factory sdpaForward(q, k, v, o, b, h, sQ, sKv, d, scale, causal), Operation fused scaled-dot-product attention (FP16, BHSD)
CuDnn has a sdpaForward method
Holds
Read Oct 3
Factory cutlassSgemm(m, n, k, alpha, A, B, beta, C), Operation FP32 SIMT GEMM
Cutlass has a cutlassSgemm method
Holds
Read Oct 3
Factory cutlassHgemm(m, n, k, alpha, A, B, beta, D), Operation FP16 tensor-core GEMM, FP32 accumulate
Cutlass has a cutlassHgemm method
Holds
Read Oct 3
Factory cutlassHgemmBatched(m, n, k, alpha, A, B, beta, C, batchCount), Operation batched FP16 tensor-core GEMM
Cutlass has a cutlassHgemmBatched method
Holds
Read Oct 3
Factory cutlassBgemm(m, n, k, alpha, A, B, beta, C), Operation BF16 tensor-core GEMM ( BFloat16Array )
Cutlass has a cutlassBgemm method
Holds
Read Oct 3
Factory cutlassGemmBiasRelu(m, n, k, A, B, bias, D), Operation fused relu(A·B + bias)
Cutlass has a cutlassGemmBiasRelu method
Holds
Read Oct 3
Factory cutlassGemmBiasGelu(m, n, k, A, B, bias, D), Operation fused gelu(A·B + bias)
Cutlass has a cutlassGemmBiasGelu method
Holds
Read Oct 3
Factory cutlassGemmBiasSilu(m, n, k, A, B, bias, D), Operation fused silu(A·B + bias)
Cutlass has a cutlassGemmBiasSilu method
Holds
Read Oct 3
Factory cutlassGemmBiasSigmoid(m, n, k, A, B, bias, D), Operation fused sigmoid(A·B + bias)
Cutlass has a cutlassGemmBiasSigmoid method
Holds
Read Oct 3
Factory cutlassGemmBiasTanh(m, n, k, A, B, bias, D), Operation fused tanh(A·B + bias)
Cutlass has a cutlassGemmBiasTanh method
Holds
Read Oct 3
Factory cutlassGemmBiasHardSwish(m, n, k, A, B, bias, D), Operation fused hardswish(A·B + bias)
Cutlass has a cutlassGemmBiasHardSwish method
Holds
Read Oct 3
Factory cusparseSpMV(rows, cols, nnz, csrRowOffsets, csrColInd, csrValues, x, y), Operation y = A·x
Cusparse has a cusparseSpMV method
Holds
Read Oct 3
Factory cusparseSpMM(rows, k, n, nnz, csrRowOffsets, csrColInd, csrValues, b, c), Operation C = A·B , dense B / C row-major
Cusparse has a cusparseSpMM method
Holds
Read Oct 3
Option withPreCompilation(), What it does JIT every task graph up front For hybrid plans this also runs each provider's prepare() , so native contexts, plans and workspaces exist before the first execution — the same reason it is a prerequisite for graph capture, Backends all
TornadoExecutionPlan has a withPreCompilation method
Holds
Read Oct 3
Option withIntraPlanConcurrency(), What it does Routes DAG-independent work (H2D, kernels, D2H) to separate role streams so it can overlap, with ordering preserved through device events derived from the bytecode dependency DAG Off by default, Backends CUDA no-op on OpenCL/Metal
TornadoExecutionPlan has a withIntraPlanConcurrency method
Holds
Read Oct 3
Option withStagedTransfers(), What it does Stages non-batched read-only transfers through pinned host memory instead of the direct path, Backends CUDA no-op on OpenCL/Metal
TornadoExecutionPlan has a withStagedTransfers method
Holds
Read Oct 3
Option withMemoryLimit("8GB"), What it does Caps device memory for the plan Worth setting explicitly when library workspaces sit alongside large operands, Backends all
TornadoExecutionPlan has a withMemoryLimit method
Holds
Read Oct 3
Option withWarmUpIterations(n), What it does Runs the whole plan n times first — transfers, compilation and execution — before the measured run, Backends all
TornadoExecutionPlan has a withWarmUpIterations method
Holds
Read Oct 3
TornadoExecutionPlan.transferToDevice(Object...) uploads the current host contents of specific objects **without running the plan**.
TornadoExecutionPlan has a transferToDevice method
Holds
Read Oct 3
…and add the same class name to src/main/resources/META-INF/services/uk.ac.manchester.tornado.runtime.library.spi.TornadoLibraryProvider.
TornadoLibraryProvider is in uk.ac.manchester.tornado.runtime.library.spi
Holds
Read Oct 3
**LibraryContext** — marker;
LibraryContext is in uk.ac.manchester.tornado.runtime.library.spi
Holds
Read Oct 3
**LibraryInvocation** — per-call payload: getArg(i), getDevicePointer(i), isReference(i), getContext(), getTuning().
LibraryInvocation is in uk.ac.manchester.tornado.runtime.library.spi
Holds
Read Oct 3
getArg(i), getDevicePointer(i), isReference(i), getContext(), getTuning().
LibraryInvocation has a getArg method
Holds
Read Oct 3
getArg(i), getDevicePointer(i), isReference(i), getContext(), getTuning().
LibraryInvocation has a getDevicePointer method
Holds
Read Oct 3
getArg(i), getDevicePointer(i), isReference(i), getContext(), getTuning().
LibraryInvocation has an isReference method
Holds
Read Oct 3
getArg(i), getDevicePointer(i), isReference(i), getContext(), getTuning().
LibraryInvocation has a getContext method
Holds
Read Oct 3
getArg(i), getDevicePointer(i), isReference(i), getContext(), getTuning().
LibraryInvocation has a getTuning method
Holds
Read Oct 3
**TornadoNativeStreamSupport** — getNativeStream(planId) / getNativeContext(planId), implemented by the CUDA backend device.
TornadoNativeStreamSupport has a getNativeStream method
Holds
Read Oct 3
**TornadoNativeStreamSupport** — getNativeStream(planId) / getNativeContext(planId), implemented by the CUDA backend device.
TornadoNativeStreamSupport has a getNativeContext method
Holds
Read Oct 3
Factory cublasSgemv, cuBLAS function cublasSgemv, Types FloatArray, Semantics y = α·op(A)·x + β·y
CuBlas has a cublasSgemv method
HoldsPR #1069 · Sep 30
Read Oct 3
Factory cublasSgemm, cuBLAS function cublasSgemm, Types FloatArray, Semantics C = α·op(A)·op(B) + β·C
CuBlas has a cublasSgemm method
HoldsPR #1069 · Sep 30
Read Oct 3
Factory cublasSgemmTF32, cuBLAS function cublasSgemm + TF32 math mode, Types FloatArray, Semantics Same as cublasSgemm , executed on TF32 Tensor Cores (~1e-4 rel error, up to 1.6x faster)
CuBlas has a cublasSgemmTF32 method
HoldsPR #1069 · Sep 30
Read Oct 3
Factory cublasSgemmStridedBatched, cuBLAS function cublasSgemmStridedBatched, Types FloatArray, Semantics C[i] = α·op(A[i])·op(B[i]) + β·C[i] for a whole batch in one call operand i lives at base + i stride in one flat array
CuBlas has a cublasSgemmStridedBatched method
HoldsPR #1069 · Sep 30
Read Oct 3
Factory cublasGemmExFP16, cuBLAS function cublasGemmEx, Types HalfFloatArray in/out, Semantics FP16 GEMM with FP32 Tensor Core accumulation
CuBlas has a cublasGemmExFP16 method
HoldsPR #1069 · Sep 30
Read Oct 3
Factory cublasGemmExFP16FP32, cuBLAS function cublasGemmEx, Types HalfFloatArray in, FloatArray out, Semantics FP16 inputs, FP32 output (standard inference config)
CuBlas has a cublasGemmExFP16FP32 method
HoldsPR #1069 · Sep 30
Read Oct 3
Factory CuBlasLt.ltMatmulFP32, cuBLAS function cublasLtMatmul, Types FloatArray, Semantics FP32 matmul (heuristic algorithm, plan cached per shape)
CuBlasLt has a ltMatmulFP32 method
HoldsPR #1069 · Sep 30
Read Oct 3
Factory CuBlasLt.ltMatmulFP16, cuBLAS function cublasLtMatmul, Types HalfFloatArray, Semantics FP16 matmul, FP32 Tensor Core accumulation
CuBlasLt has a ltMatmulFP16 method
HoldsPR #1069 · Sep 30
Read Oct 3
Factory CuBlasLt.ltMatmulBiasFP16, cuBLAS function cublasLtMatmul + BIAS, Types HalfFloatArray, Semantics C = op(A)·op(B) + bias, fused
CuBlasLt has a ltMatmulBiasFP16 method
HoldsPR #1069 · Sep 30
Read Oct 3
Factory CuBlasLt.ltMatmulGeluBiasFP16, cuBLAS function cublasLtMatmul + GELU_BIAS, Types HalfFloatArray, Semantics C = GELU(op(A)·op(B) + bias), fully fused transformer MLP block
CuBlasLt has a ltMatmulGeluBiasFP16 method
HoldsPR #1069 · Sep 30
Read Oct 3
**CUDA Graphs:** supported — plans (whose creation allocates a device work area) are built in the provider's prepare() hook, which the runtime invokes before capture starts, so the FFT pipeline is captured and replayed like any other library task (see TestCuFft#testRoundTripWithCudaGraph).
CuFftLibraryProvider has a prepare method
Holds
Read Oct 3
Factory Cusparse.cusparseSpMV(rows, cols, nnz, csrRowOffsets, csrColInd, csrValues, x, y), Operation y = A · x
Cusparse has a cusparseSpMV method
Holds
Read Oct 3
Factory Cusparse.cusparseSpMM(rows, k, n, nnz, csrRowOffsets, csrColInd, csrValues, B, C), Operation C = A · B (dense B k×n, C rows×n)
Cusparse has a cusparseSpMM method
Holds
Read Oct 3
those have no C ABI to call, so here they are Java records in registries, and the long handles the backend passes around for programs, kernels and events are registry keys rather than raw pointers (real Metal objects -- devices, queues, buffers, libraries, pipelines, command buffers -- stay raw pointers).
Metal depends on MetalObjects
HoldsPR #1069 · Sep 30
Read Oct 3
The backend classes (Metal, MetalPlatform, MetalDevice, MetalContext, MetalCommandQueue, MetalProgram, MetalKernel, MetalEvent, NativeCommandQueue) delegate to it in place of their old native methods.
MetalPlatform depends on MetalObjects
HoldsPR #1069 · Sep 30
Read Oct 3
The backend classes (Metal, MetalPlatform, MetalDevice, MetalContext, MetalCommandQueue, MetalProgram, MetalKernel, MetalEvent, NativeCommandQueue) delegate to it in place of their old native methods.
MetalDevice depends on MetalObjects
HoldsPR #1069 · Sep 30
Read Oct 3
The backend classes (Metal, MetalPlatform, MetalDevice, MetalContext, MetalCommandQueue, MetalProgram, MetalKernel, MetalEvent, NativeCommandQueue) delegate to it in place of their old native methods.
MetalContext depends on MetalObjects
HoldsPR #1069 · Sep 30
Read Oct 3
The backend classes (Metal, MetalPlatform, MetalDevice, MetalContext, MetalCommandQueue, MetalProgram, MetalKernel, MetalEvent, NativeCommandQueue) delegate to it in place of their old native methods.
MetalCommandQueue depends on MetalObjects
HoldsPR #1069 · Sep 30
Read Oct 3
The backend classes (Metal, MetalPlatform, MetalDevice, MetalContext, MetalCommandQueue, MetalProgram, MetalKernel, MetalEvent, NativeCommandQueue) delegate to it in place of their old native methods.
MetalProgram depends on MetalObjects
HoldsPR #1069 · Sep 30
Read Oct 3
The backend classes (Metal, MetalPlatform, MetalDevice, MetalContext, MetalCommandQueue, MetalProgram, MetalKernel, MetalEvent, NativeCommandQueue) delegate to it in place of their old native methods.
MetalKernel depends on MetalObjects
HoldsPR #1069 · Sep 30
Read Oct 3
The backend classes (Metal, MetalPlatform, MetalDevice, MetalContext, MetalCommandQueue, MetalProgram, MetalKernel, MetalEvent, NativeCommandQueue) delegate to it in place of their old native methods.
MetalEvent depends on MetalObjects
HoldsPR #1069 · Sep 30
139 rules from 9 docs

See these where the change is. The browser extension puts a pull request's doc findings, and a diagram of what it changed, on the GitHub page itself, so the code and what your docs say about it are side by side.