Conversation
LLVM replaced the single BranchInst class with separate UncondBrInst and CondBrInst classes and removed BranchInst entirely (llvm-project 64fc793dd100, 5b4015e55961, 464639bdcc8c). That landed mid-cycle during LLVM 24 development, so it cannot be keyed off LLVM_VERSION_MAJOR; CMake probes for llvm::CondBrInst and defines ENZYME_SPLIT_BRANCH_INST. The new classes share no base other than Instruction, so the shims in BranchCompat.h take and return Instruction*, which already provides getNumSuccessors/getSuccessor/setSuccessor on every supported LLVM. Only the condition accessors and branch creation need to know which class they are dealing with, and those are confined to the header. Two other in-flight LLVM changes come along for the ride: * InsertPosition no longer constructs from Instruction*, so the one CallInst::Create that relied on it passes an iterator (LLVM 19+). * ScalarEvolution::getLosslessPtrToIntExpr was renamed to getPtrToAddrExpr, probed for as ENZYME_HAS_SCEV_PTRTOADDR. Verified against LLVM 22 (pre-split) and LLVM 24 (post-split). On LLVM 22 check-enzyme reports an identical set of results before and after this commit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
LLVM removed the deprecated BasicBlock::getFirstNonPHI (llvm-project 62c5ede9fd14), leaving only the iterator-flavoured getFirstNonPHIIt. Add a getFirstNonPHI(BasicBlock*) shim alongside the existing getFirstNonPHIOrDbg wrappers and move the 19 call sites onto it. The iterator overload has been available since LLVM 18, so older LLVMs keep the old spelling. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
flang lowers a type-bound procedure call to a load out of the derived
type's binding table followed by an indirect call, so the bound procedures
have no direct call site at all -- their address only ever appears as a
`ptrtoint` entry inside the table. Enzyme reaches them from
GetOrCreateShadowConstant walking that table, and GetOrCreateShadowFunction
then seeded every pointer argument with an empty TypeTree, leaving
TypeAnalysis to recover the layout of each argument from the body alone.
Two additions, both in terms of a new pair of utilities in Utils.cpp:
* getDevirtualizedCallee constant-folds an indirect callee expression
through loads out of provably immutable memory (LLVM's
ConstantFoldLoadFromConstPtr only folds constant globals with a
definitive initializer), so a non-null answer is guaranteed. TypeAnalyzer
::visitCallBase uses it so that visitIPOCall applies to such calls.
* getIndirectCallCandidates over-approximates the target set of a
`call (inttoptr (load (base + const offset)))` by collecting, out of
every constant dispatch-table global in the module, the function at that
same offset -- the base itself need not be known, which is what makes it
apply to a dispatch off a runtime class descriptor.
GetOrCreateShadowFunction uses it to find the call sites that may reach
the function it is about to differentiate and takes the *meet* of their
argument types: an over-approximated call site only weakens the result,
and if any matching-signature indirect call cannot be bounded we fall
back to no information at all.
Enzyme-created shadow globals are tagged with "enzyme_shadow_of" so that a
shadow table is not mistaken for a source-level one. The whole thing can be
turned off with -enzyme-devirtualize=0.
On a whole-program flang module containing DVODE this takes the
"cannot handle unknown binary operator" count from 27 to 19.
visitGEPOperator only typed a GEP's indices as Integer when the GEP was `inbounds`. flang never emits `inbounds` for array element addresses -- it emits `nusw nuw`, from XArrayCoorOp lowering in flang's CodeGen.cpp -- so a Fortran array index was never typed. That is self-reinforcing: the index is not Integer because the GEP is not inbounds and the base is not yet a known Pointer, and the base never becomes a Pointer because pointer propagation requires either inbounds or indices that are already integral. Both the index chain and the array come out completely untyped. Activity analysis then cannot prove the arithmetic inactive, and integer address computations reach visitBinaryOperator's unhandled case as "cannot handle unknown binary operator". `nusw` is the premise the rule actually needs: the index is added as a signed byte offset which does not wrap, so it is an offset rather than something that might itself be a pointer. `inbounds` additionally guarantees the result stays within the object, which this rule does not rely on. The pointer propagation below is deliberately left keyed on isInBounds(); once the indices are typed Integer, its existing allIntegral path enables propagation on the next fixpoint iteration without weakening that stronger premise. Two tests: TypeAnalysis/gepnusw.ll pins the rule on hand-written IR, with an inbounds twin that must analyse identically, and Fortran/TypeAnalysis/gep_nusw.f90 shows the same shape is what flang actually produces. Both fail without the change. On a whole-program Fortran module differentiating DVODE this takes "cannot handle unknown binary operator" from 27 to 17, eliminating every failure in dvindy and dvjust; the remainder are a separate, unrelated cause. check-enzyme and check-typeanalysis show no newly failing tests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
invertPointerM's null-shadow path stores the value into an alloca so it can zero the float lanes, and reaches that store via getNewFromOriginal. Only instructions and arguments live in originalToNewFn; constants are shared between the old and new function and map to themselves, so a constant operand asserted with `f != originalToNewFn.end()`. Reachable from Fortran: flang's derived type descriptors are constant structs holding ptrtoint of globals, and differentiating a store of one recurses through invertPointerM into exactly such an operand. This is a partial fix. It removes the assert, after which the same input hits a second, separate limitation one level up: the constant-aggregate recursion requires each element's shadow to be a Constant, while the null-shadow path emits alloca/store/load. Masking a constant would have to be constant-folded rather than emitted as instructions. Left alone here. No test: the reduced cases I could construct do not reach this path, and the whole-program input that does still crashes for the reason above. check-enzyme and check-typeanalysis are unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
invertPointerM masks a constant's float lanes by storing it into an alloca, zeroing those bytes and loading it back. That cannot be used for an element of a constant aggregate, whose shadow must itself be a constant: the recursion does cast<Constant> on the result and aborts. Fold the mask instead when the value is a constant. maskFloatsInConstant walks the constant structurally against the TypeTree, zeroing floating point leaves and carrying everything else through, and returns null when the result cannot be expressed as a constant -- a scalar only partly covered by float -- in which case the caller still emits the runtime masking. The common case is cheaper than it looks. anyFloat treats an unknown byte as possibly float, so a constant TypeAnalysis knows nothing about takes this path with no byte actually known to be float; the byte loop would zero nothing, and the fold returns the constant itself. Together with the previous commit this removes the abort on a whole-program Fortran module differentiating DVODE: it now reports its errors and exits cleanly rather than crashing partway through. The error count rises 82 -> 135 purely because the pass now runs to completion; the categories are unchanged (mismatched activity, cannot deduce type of copy, no derivative found). No test: reaching this path needs a constant aggregate whose element has an unknown sub-TypeTree while the aggregate has float elsewhere, which the reduced cases I built do not produce -- they type the element as Pointer and take an earlier path. check-enzyme and check-typeanalysis are unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
considerRustDebugInfo scans for DbgDeclareInst. LLVM 19 moved debug info from intrinsics to records, so on any newer LLVM that scan matches nothing and -enzyme-rust-type is silently inert -- a module with 288 #dbg_declare records yields zero hits. Walk the record range as well. Untested against Rust: every rust*.ll test in the suite already fails on this LLVM for unrelated reasons (typed-pointer CHECK lines), so there is nothing here that can exercise the parser end to end. The change is off by default, and check-enzyme and check-typeanalysis are unchanged. Found while checking whether debug info could supply the types flang does not. It cannot, at least not as the parser stands: with this fix and the flag on, a whole-program Fortran module goes from 135 errors to 134, and the 17 untyped integer counters are untouched. The parser only types allocas reached from a declare, not the pointer arguments the solver state arrives in, and it asserts on Fortran DWARF tags it does not model. Left as is rather than degraded, since that assert is a deliberate signal. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
KnownInactiveFunctionsStartingWith already lists "f90io", classic flang's I/O runtime. LLVM flang names the same thing _FortranAio*, which was not covered, so differentiating any Fortran routine that prints -- DVODE's xerrwd, for instance -- failed with "No forward mode derivative found". Also add _FortranAClassIs and _FortranATrim, a type query and character handling; neither moves floating point data. _FortranAAssign is deliberately not included. It copies data which may include reals, so declaring it inactive would silently drop a derivative; it needs a rule that assigns the shadow too. On a whole-program Fortran module differentiating DVODE, with loose types, this takes the error count from 17 to 1, the remainder being that _FortranAAssign. check-enzyme and check-typeanalysis are unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
_FortranAAssign(to, from, sourceFile, sourceLine) is flang's assignment between descriptors. It moves data which may include reals, so it cannot be declared inactive; its derivative is the same assignment applied to the shadows. Passing the shadows straight to the runtime does not work. A shadow is only shadow storage: element length, rank, type and the derived type addendum are never written to it, because none of that is differentiated, so the runtime reads a zeroed rank and walks off the end. That segfaults inside DerivedAssignTicket::Begin. Copy the primal descriptor over the shadow first, keeping the shadow's own base_addr, so the runtime is handed a well formed descriptor addressing the shadow data. The size comes from the underlying alloca, which is how flang emits these; if it cannot be determined the rule declines to handle the call and it is reported as underivable rather than guessed at. What this does not do: make Fortran derived type assignment differentiate correctly end to end. On the DVODE reproducer the program now links and runs to completion where it previously crashed, but the result is still wrong. That is upstream of this rule -- disabling just the shadow call leaves the primal trace equally corrupt (garbage rather than zeros), so the generated code is being broken by the surrounding type ambiguity that -enzyme-loose-types papers over, not by the assignment. Related: #2963, which scaffolds Fortran intrinsic dispatch but leaves descriptor handling as a TODO. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
llvm.fake.use keeps a value live for the debugger and computes nothing, so there is nothing to differentiate, but it reached the unknown-intrinsic paths and stopped differentiation outright. flang emits it for every local at -O0 -g, where DVODE's dvhin alone has 21 of them; a whole-program forward-mode build failed with 20 "cannot handle (forward) unknown intrinsic" errors and now succeeds with none. Handled in both switches. The forward mode one is the path that fired here; the other is the same argument. This matters beyond -O0: it is what lets the reproducer be built without inlining, which is how the remaining wrong answer was pinned down as a deterministic miscompilation rather than uninitialised memory. check-enzyme and check-typeanalysis are unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
TypeAnalysis turns the bytes of a constant data array into Anything (ReplaceIntWithAnything, for anything at least 16 bits wide), and visitMemTransferCommon purges Anything out of the source query. For a copy whose source is a character literal that leaves nothing at all, and the copy is reported as undeducible. A copy out of a constant whose type contains no floating point member anywhere cannot carry a derivative, so it is integer data. Checking the type rather than the value keeps it conservative: a derived type initializer holding reals is excluded and still gets analysed properly. This is what flang's error message strings hit. On the DVODE reproducer it removes the last "Cannot deduce type of copy" failures, which together with -enzyme-max-type-offset and runtime activity is what allows that program to be built with real type analysis rather than -enzyme-loose-types. Also guard the _FortranAAssign rule against a shadow that is the primal itself, where assigning through it would overwrite primal data. check-enzyme and check-typeanalysis are unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
GetOrCreateShadowFunction builds an FnTypeInfo from scratch, so a pointer argument arrives at TypeAnalysis carrying nothing. For a function whose only references are dispatch table entries there is no call site to recover the layout from either, which leaves the body as the sole source of type information -- and for a solver's integer state there is nothing in the body to recover it from. An explicit enzyme_type parameter attribute is exactly the missing statement, and every other consumer already honors it (Enzyme.cpp for the differentiated function, TypeAnalysis.cpp for call sites). Use it here too, in preference to the type-derived default. With it, annotating DVODE's type-bound procedures with their descriptor layout takes "cannot handle unknown binary operator" on a whole-program Fortran module from 19 to 0, and the module differentiates with real type analysis rather than -enzyme-loose-types. Note the annotation must be paired with an -enzyme-max-type-offset large enough to cover the struct, or the tree is silently truncated at 500 and the fields past it stay untyped. check-enzyme and check-typeanalysis are unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An integer operation narrower than a pointer cannot be pointer arithmetic, and integer arithmetic on the bits of a float is not meaningful, so an operation of unknown type here is an integer one and its result is constant. Apply that to the integer binary operators and to the integer intrinsics (smax, umax, abs, the overflow intrinsics) in forward mode. flang needs this. The integer counters of a derived type -- DVODE's nfe, nje and nst in dvode_data_t -- are only ever loaded, incremented and stored, so nothing ever gives them a type, and forward mode then has to differentiate an integer add. On the DVODE reproducer that was 17 fatal errors from the binary operators and one more from llvm.smax. The only way through was -enzyme-loose-types, which guesses a type everywhere rather than just where the guess is sound; the reproducer now links with no type-analysis flag. This does not change any derivative the reproducer produces: with and without -enzyme-loose-types the results were bit-for-bit identical, so the flag was only ever the reason the build needed a global guess. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011Z6KU642kqfZcHh6TAExv3
Collaborator
|
@vchuravy something I forgot to mention in the meeting yesterday: we can pick up some of your fixes here and move towards getting them merged if you like? |
Member
Author
|
Please do! |
This was referenced Sep 8, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.