[pull] main from llvm:main - #1707
Merged
Merged
Conversation
Regular cross-team reductions have two phases: the intra-team reduction and the inter-team reduction. Atomic cross-team reductions replace the second phase with a atomic instruction which is used by the main thread of each team to directly fold the result of the intra-team reduction into the final result. Since this requires a combination of "data type" and "combine operation" for which an atomic instruction is available, only some (but very common) reductions can be transformed to atomic reductions. In cases where multiple reductions are performed on the same construct, the atomic path is only taken if all reductions can be transformed. Otherwise, we fall back to the regular cross-team reduction using a buffer with per-team slots. This is not strictly necessary, but hybrid reductions would induce more complexity with questionable benefit. Selecting an atomic path might not be the best option for every situation, which is why it is not enabled by default. Instead, it can be enabled via `-fopenmp-target-atomic-reduction`. Note that enabling the atomic path will not *force* atomic reductions. They will only be applied if possible, as described above. The performance (measured with https://github.com/ro-i/xteam-test @ c71339705091500f731e2a39f247d2660bacbdce, array size 177,777,777) is up to +15% faster (aka, more throughput) for supported reductions on a gfx942, with no noticeable regressions. Example: - sum reduction, type double: +10.22% faster - sum reduction, type uint: +15.57% faster - sum reduction, type ulong: +13.31% faster On a gfx90a, there is little to negative benefit: - sum reduction, type double: -4.32% faster (aka, slower) - sum reduction, type uint: +3.08% faster - sum reduction, type ulong: +1.68% faster Claude assisted with this patch.
An RTBridge Caller is a controller-side handle for calling a function in the runtime. Until now the abstraction assumed every such function was a trampoline -- a runtime function whose job is to invoke *another* function at an address the controller supplies (run-as-main, run-as-int, etc.) -- so every Caller carried a dedicated ExecutorAddr parameter for that target. Generalize Callers to call runtime functions of any shape. Invoking a supplied target is now just one kind of call, with the target address an ordinary leading argument rather than a built-in parameter: e.g. MainCaller becomes Caller<int64_t(ExecutorAddr, ArrayRef<std::string>)>. The SPS signatures already led with an SPSExecutorAddr for the target, so this is a pure interface change -- the SPS wrappers and all call sites are unaffected. It lets Callers model runtime functions that do the work themselves, such as the memory-access wrappers, rather than only those that dispatch to another function.
`clang++` defaults to `-stdlib=libc++` on NetBSD. When building with
both `clang` and `libcxx` included, the freshly built `clang++` fails to
find `<__config_site>`:
```
In file included from /usr/include/strings.h:68:
In file included from bin/../include/c++/v1/string.h:57:
bin/../include/c++/v1/__config:13:10: fatal
error: '__config_site' file not found
13 | #include <__config_site>
| ^~~~~~~~~~~~~~~
```
The file is present in `include/<triplet>/c++/v1`, but that isn't
searched by default. NetBSD has its own version of addLibCxxIncludePaths
which misses that directory.
This patch removes `NetBSD::addLibCxxIncludePaths` in favour of the
generic version in `Gnu.cpp`. The current code also adds
`/usr/include/c++`, although this directory only contains empty
directories in a default installation. It is only used when a bundled
version of LLVM is installed, which is not usually the case, and even
then contains a static version of `__config_site` that only applies to
`libcxxrt`.
Tested on `amd64-pc-netbsd10.1`, `x86_64-pc-solaris2.11`,
`x86_64-pc-linux-gnu`, and `x86_64-pc-freebsd15.1`.
On Windows/COFF, a dllimport call is emitted as an indirect call through an `__imp_` IAT slot (`callq *__imp_bar(%rip)`), and even a direct call to a library function is expected to bind to a thunk supplied by an import library. Today a JIT client must produce those import libraries themselves. `AutoImportGenerator` synthesizes them on demand instead. Bound to a single dynamic library via `AutoImportGenerator::Load(ES, ObjLinkingLayer, "/path/to/lib.dll")`. For each referenced export `X`, lazily synthesizes an `__imp_X` pointer slot holding `X`'s address in the library plus an `X` thunk that jumps through it, so both `__imp_`-mediated and direct references resolve. The library's export table is the authority: a name the library does not export is left unresolved, so the link fails exactly as a static link against the corresponding import library would (no silent invention of symbols). All synthesized stubs are owned by a single `ResourceTracker` (`getImportStubsResourceTracker()`), so a client can reclaim every synthesized slot/thunk in one step; subsequent imports start a fresh tracker. Relationship to `DLLImportDefinitionGenerator`: that generator resolves the underlying symbol through the JITDylib's link order. `AutoImportGenerator` is bound to one specific library and treats its export table as authoritative, giving fail-as-static-linker semantics. x86_64 and in-process execution only. - "Easy mode": every import is assumed to be a function; code and data are not distinguished, so data imports are unsupported (clients with data imports must supply an import library or use `__declspec(dllimport)`). `&X` resolves to the synthesized thunk, not the implementation in the target library. `llvm-jitlink -auto-import=<lib>` flag to attach the generator (in-process; errors if combined with out-of-process execution). Documentation in `llvm/docs/ORCv2.rst`. Partly implements github issue: #190122 In the comment section of the github issue there is this comment #190122 (comment) This PR implements point 3 ("Easy mode" generator)
…, NFC Reviewers: Pull Request: #213538
The change is made in HexagonExpandCondsets::(predicate) function. The debug instructions are not predicable as they cannot be separated into conditional branches. So while predicating instructions in a machine basic block if we encounter any debug instructions we need to skip these instructions and continue with other instructions. The scan that collects the registers defined and used between the definition of the source register and the conditional transfer bailed out as soon as it saw a non-virtual register operand, which a DBG_VALUE can have. The transfer was then left as an unconditional A2_asrh plus an A2_tfrf instead of being folded into a single predicated A4_pasrhf, so again the generated code differed depending on whether debug info was enabled. Co-authored-by: Chandana Sinderikeri <csinderi@qti.qualcomm.com>
…2822) Do not call the Sema actions for absent, contains, and nullary assumption clauses after the parser has diagnosed that the clause is not allowed on the current directive. Add assertions documenting that these Sema actions must only receive clauses allowed on the current directive, and add tests covering all affected clause kinds. Fixes #212780.
- Add LLVM_LIBC_ADD_FUNCTION_C_ALIAS macro to add another C alias public symbol to a function. - Add LIBC_CONF_SCANF_PROVIDE_ISOC99_ALIASES config - Add __isoc99_fscanf for generic fscanf target if LIBC_CONF_SCANF_PROVIDE_ISOC99_ALIASES is set. - Similarly: scanf, vfscanf, vscanf.
Found when looking through the generated Python file. `sig` doesn't exist in that context, it should be `idx`.
is_thread_crashed maps how each platform reports the bad access the test suite uses to simulate a crash, and had no case for Wasm, where there are no signals and a bad access raises a trap that a runtime reports as an exception. Without one it fell through to a description match that never held, so a crashed thread read as running fine.
This avoids defining an `enum`, which appears to be expensive for Clang.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
See Commits and Changes for more details.
Created by
pull[bot] (v2.0.0-alpha.4)
Can you help keep this open source service alive? 💖 Please sponsor : )