clang-p2996

Author	SHA1	Message	Date
Ross Brunton	02d2a1646a	[Offload] Fix entry_points.td test (#145292 ) This was broken as part of #144494 , and just needs an update to the check lines.	2025-06-23 11:09:08 +01:00
Ross Brunton	613c38a992	[Offload] Fix type mismatch warning in test (#143700 )	2025-06-23 10:14:12 +01:00
Joseph Huber	3f1de197b1	[Offload] Rework compiling device code for unit test suites (#144776 ) Summary: I'll probably want to use this as a more generic utility in the future. This patch reworks it to make it a top level function. I also tried to decouple this from the OpenMP utilities to make that easier in the future. Instead, I just use `-march=native` functionality which is the same thing. Needed a small hack to skip the linker stage for checking if that works. This should still create the same output as far as I'm aware.	2025-06-20 10:31:54 -05:00
Ross Brunton	f242360e15	[Offload] Add type information to device info nodes (#144535 ) Rather than being "stringly typed", store values as a std::variant that can hold various types. This means that liboffload doesn't have to do any string parsing for integer/bool device info keys.	2025-06-20 09:05:05 -05:00
Ross Brunton	e0633d59b9	[Offload] Check for initialization (#144370 ) All entry points (except olInit) now check that offload has been initialized. If not, a new `OL_ERRC_UNINITIALIZED` error is returned.	2025-06-20 09:04:50 -05:00
Ross Brunton	53336ad488	[Offload] Move (most) global state to an `OffloadContext` struct (#144494 ) Rather than having a number of static local variables, we now use a single `OffloadContext` struct to store global state. This is initialised by `olInit`, but is never deleted (de-initialization of Offload isn't yet implemented). The error reporting mechanism has not been moved to the struct, since that's going to cause issues with teardown (error messages must outlive liboffload).	2025-06-19 16:02:03 -05:00
Jan Patrick Lehr	dd65e6e060	[Offload][libc] Add cmake cache AMDGPU buildbot (#144500 ) An upcoming libc4GPU buildbot will be using this CMake cache file for its build configuration.	2025-06-17 20:51:40 +02:00
Ross Brunton	e6a3579653	[Offload] Replace device info queue with a tree (#144050 ) Previously, device info was returned as a queue with each element having a "Level" field indicating its nesting level. This replaces this queue with a more traditional tree-like structure. This should not result in a change to the output of `llvm-offload-device-info`.	2025-06-13 09:22:47 -05:00
Ethan Luis McDonough	daee5eee85	[Offload][PGO] Fix new GPU PGO tests (#143645 ) `pgo_atomic_teams.c` and `pgo_atomic_threads.c` currently are set to run on NVPTX despite the changes for that target not being upstreamed yet. This patch also replaces instances of `llvm-profdata` with `%profdata` in those tests.	2025-06-12 11:14:21 -05:00
Ross Brunton	4f60321ca1	[Offload] Add `ol_dimensions_t` and convert ranges from size_t -> uint32_t (#143901 ) This is a three element x, y, z size_t vector that can be used any place where a 3D vector is required. This ensures that all vectors across liboffload are the same and don't require any resizing/reordering dances.	2025-06-12 09:59:59 -05:00
Abhinav Gaba	02b6849cf1	[Clang][OpenMP] Fix mapping of arrays of structs with members with mappers (#142511 ) This builds upon #101101 from @jyu2-git, which used compiler-generated mappers when mapping an array-section of structs with members that have user-defined default mappers. Now we do the same when mapping arrays of structs.	2025-06-11 19:03:55 +00:00
Kewen12	bbe59e19b6	[OpenMP][Offload] Update the Logic for Configuring Auto Zero-Copy (#143638 ) Summary: Currently the Auto Zero-Copy is enabled by checking every initialized device to ensure that no dGPU is attached to an APU. However, an APU is designed to comprise a homogeneous set of GPUs, therefore, it should be sufficient to check any device for configuring Auto Zero-Copy. In this PR, it checks the first initialized device in the list. The changes in this PR are to clearly reflect the design and logic of enabling the feature for further improving the readibility.	2025-06-11 14:12:54 -04:00
Ethan Luis McDonough	67ff66e677	[PGO][Offload] Fix offload coverage mapping (#143490 ) This pull request fixes coverage mapping on GPU targets. - It adds an address space cast to the coverage mapping generation pass. - It reads the profiled function names from the ELF directly. Reading it from public globals was causing issues in cases where multiple device-code object files are linked together.	2025-06-10 20:19:38 -05:00
Ross Brunton	637df705e5	[Offload] Add `OFFLOAD_INCLUDE_TESTS` (#143388 ) This is a cmake variable which, if set to `OFF` will disable building of tests. It defaults to the value of `LLVM_INCLUDE_TESTS`.	2025-06-09 10:27:40 -05:00
Callum Fare	835497a4dc	[Offload] Make olMemcpy src parameter const (#143161 )	2025-06-06 10:25:00 -05:00
Ross Brunton	269c29ae67	[Offload] Allow setting null arguments in olLaunchKernel (#141958 )	2025-06-06 07:05:11 -05:00
Joseph Huber	051945304b	[Offload] Fix APU detection for MI300 testing (#143026 ) Summary: We have this check when the target is MI300 but it fails if this environment variable isn't set. Set a default value of '0' if not present so that will be converted to bool false.	2025-06-05 15:31:55 -05:00
Callum Fare	f44df93a9c	[Offload] Explicitly create directories that contain tablegen output (#142817 ) This isn't required when building with Ninja, but with the Makefile generator these directories don't get implicitly created.	2025-06-04 13:46:19 -05:00
Callum Fare	817af2ddf2	[Offload] Fix missing dependencies in Offload API generation (#142776 ) Thanks to @RossBrunton for spotting this. We attempt to clang-format the generated Offload header files, but if clang-format isn't available we just copy the generated files instead. That fallback path was missing the correct dependencies. Fixes #142756	2025-06-04 08:51:50 -05:00
Callum Fare	b78bc35d16	[Offload] Don't check in generated files (#141982 ) Previously we decided to check in files that we generate with tablegen. The justification at the time was that it helped reviewers unfamiliar with `offload-tblgen` see the actual changes to the headers in PRs. After trying it for a while, it's ended up causing some headaches and is also not how tablegen is used elsewhere in LLVM. This changes our use of tablegen to be more conventional. Where possible, files are still clang-formatted, but this is no longer a hard requirement. Because `OffloadErrcodes.inc` is shared with libomptarget it now gets generated in a more appropriate place.	2025-06-03 10:39:04 -05:00
Jan Patrick Lehr	e97f42e931	[OpenMP][Offload] Fix typo in error message (#142589 ) It appears that the spelling was incorrect in those test cases. At least on machines with ROCm version > 6.3. I had no chance to test with ROCm version version < 6.2 and would be interested in the result if someone has the chance.	2025-06-03 07:33:45 -05:00
Joseph Huber	eb9ed93fce	[Offload] Optimistically accept SM architectures (#142399 ) Summary: We try to clamp these to ones known to work, but we should probably just optimistically accept these. I'd prefer to update the flag check, but since NVIDIA refuses to publish their ELF format it's too much effort to reverse engineer. Fixes: https://github.com/llvm/llvm-project/issues/138532	2025-06-02 14:32:05 -05:00
Ross Brunton	e83c80340f	[Offload] Split offload unittests into multiple files (#142418 ) Rather than a single `offload.unittests` file, this will produce `device.unittests`, `event.unittests`, etc.. This should reduce time spent building tests, and make it easier to manually run a subset of the tests. Note that `check-offload-unit` will still run all the tests.	2025-06-02 11:48:12 -05:00
Joseph Huber	5b8031a7f7	[Offload][AMDGPU] Correctly handle variable implicit argument sizes (#142199 ) Summary: The size of the implicit argument struct can vary depending on optimizations, it is not always the size as listed by the full struct. Additionally, the implicit arguments are always aligned on a pointer boundary. This patch updates the handling to use the correctly aligned offset and only initialize the members if they are contained in the reported size. Additionally, we modify the `alloc` and `free` routines to allow `alloc(0)` and `free(nullptr)` as these are mandated by the C standard and allow us to easily handle cases where the user calls a kernel with no arguments.	2025-06-02 09:35:16 -05:00
Ross Brunton	41e22aa31b	[Offload] Set size correctly in olLaunchKernel cts test (#142398 ) It was previously not scaled by `sizeof(uint32_t)`.	2025-06-02 09:27:09 -05:00
Joseph Huber	b26baf1779	[Offload] Make AMDGPU plugin handle empty allocation properly (#142383 ) Summary: `malloc(0)` and `free(nullptr)` are both defined by the standard but we current trigger erros and assertions on them. Fix that so this works with empty arguments.	2025-06-02 08:12:20 -05:00
Ross Brunton	7efb79b705	[Offload] Fix Error checking (#141939 ) All errors must be checked - this includes the local variable we were using to increase the lifetime of `Res`. As we were not explicitly checking it, it resulted in an `abort` in debug builds.	2025-05-29 08:17:08 -05:00
Ross Brunton	a1191b4875	[Offload] Fix broken tablegen test after #140879 (#141796 )	2025-05-28 11:30:15 -05:00
Joseph Huber	0ebe5557d9	[Offload] Add specifier for the host type (#141635 ) Summary: We use this sepcial type to indicate a host value, this will be refined later but for now it's used as a stand-in device for transfers and queues. It needs a special kind because it is not a device target as the other ones so we need to differentiate it between a CPU and GPU type. Fixes: https://github.com/llvm/llvm-project/issues/141436	2025-05-28 08:51:14 -05:00
Joseph Huber	a9b64bb318	[Offload] Fix segfault when looking for host device name (#141632 ) Summary: This is done using the generic device into pointe, but no such thing exists for the host device, leading to a segfault. This patch fixes that for now, but in the future we should probably be more careful in general handling the possibility that the handle is null everywhere. Fixes: https://github.com/llvm/llvm-project/issues/141434	2025-05-27 13:43:29 -05:00
Joseph Huber	20f9f1fc02	[Offload][NFCI] Remove coupling to `omp` target for version scripting (#141637 ) Summary: This is a weird dependency on libomp just for testing if version scripts work. We shouldn't need to do this because LLVM already checks for this. I believe this should be available as well in standalone when we call `addLLVM` but I did not test that directly.	2025-05-27 13:43:07 -05:00
Ross Brunton	7e9d708be0	[Offload] Use llvm::Error throughout liboffload internals (#140879 ) This removes the `ol_impl_result_t` helper class, replacing it with `llvm::Error`. In addition, some internal functions that returned `ol_errc_t` now return `llvm::Error` (with a fancy message).	2025-05-27 13:42:56 -05:00
Johannes Doerfert	57a90edacd	[OpenMP][GPU][FIX] Enable generic barriers in single threaded contexts (#140786 ) The generic GPU barrier implementation checked if it was the main thread in generic mode to identify single threaded regions. This doesn't work since inside of a non-active (=sequential) parallel, that thread becomes the main thread of a team, and is not the main thread in generic mode. At least that is the implementation of the APIs today. To identify single threaded regions we now check the team size explicitly. This exposed three other issues; one is, for now, expected and not a bug, the second one is a bug and has a FIXME in the single_threaded_for_barrier_hang_1.c file, and the final one is also benign as described in the end. The non-bug issue comes up if we ever initialize a thread state. Afterwards we will never run any region in parallel. This is a little conservative, but I guess thread states are really bad for performance anyway. The bug comes up if we optimize single_threaded_for_barrier_hang_1 and execute it in Generic-SPMD mode. For some reason we loose all the updates to b. This looks very much like a compiler bug, but could also be another logic issue in the runtime. Needs to be investigated. Issue number 3 comes up if we have nested parallels inside of a target region. The clang SPMD-check logic gets confused, determines SPMD (which is fine) but picks an unreasonable thread count. This is all benign, I think, just weird: ``` #pragma omp target teams #pragma omp parallel num_threads(64) #pragma omp parallel num_threads(10) {} ``` Was launched with 10 threads, not 64.	2025-05-20 19:33:54 -07:00
Ross Brunton	c19a3cb613	[Offload] Make OffloadAPI gtest error messages more readable (#140728 )	2025-05-20 08:50:26 -05:00
Ross Brunton	050892d2f8	[Offload] Use new error code handling mechanism and lower-case messages (#139275 ) [Offload] Use new error code handling mechanism This removes the old ErrorCode-less error method and requires every user to provide a concrete error code. All calls have been updated. In addition, for consistency with error messages elsewhere in LLVM, all messages have been made to start lower case.	2025-05-20 08:50:20 -05:00
Ross Brunton	1532ee6916	[Offload] Add Error Codes to PluginInterface (#138258 ) A new ErrorCode enumeration is present in PluginInterface which can be used when returning an llvm::Error from offload and PluginInterface functions. This enum must be kept up to sync with liboffload's ol_errc_t enum, so both are automatically generated from liboffload's enum definition. Some error codes have also been shuffled around to allow for future work. Note that this patch only adds the machinery; actual error codes will be added in a future patch. ~~Depends on #137339 , please ignore first commit of this MR.~~ This has been merged.	2025-05-19 09:38:34 -05:00
Ethan Luis McDonough	1043810769	[PGO][Offload] Update PGO GPU tests (#132262 )	2025-05-14 17:17:52 -05:00
Dhruva Chakrabarti	f965996cfb	[Offload] Remove unused field IsBareKernel. (#139815 )	2025-05-13 17:35:55 -07:00
agozillon	f687ed9ff7	[Flang][OpenMP] Initial defaultmap implementation (#135226 ) This aims to implement most of the initial arguments for defaultmap aside from firstprivate and none, and some of the more recent OpenMP 6 additions which will come in subsequent updates (with the OpenMP 6 variants needing parsing/semantic support first).	2025-05-12 16:30:43 +02:00
Joseph Huber	d60eeda2e5	[Offload] Do not load images from the same descriptor on the same device (#139147 ) Summary: Right now we generally assume that we have one image per device. The binary descriptor represents a single 'compilation'. This means that each image is going to contain the same code built for different architectures when used through the OpenMP interface. This is problematic when we have cases where the same code will then be loaded multiple times (like wiht sm_80, sm_89 or the generic GFX ISAs). This patch is the quick and dirty slution, we just prevent this from happening at all. This means we use the first one we find, which might not be overly optimal, but it should be better than the alternative. Note that this does not affect shared library loads as it is per binary descriptor, not per device.	2025-05-09 08:21:40 -05:00
agozillon	b291cfcad4	[Flang][OpenMP] Generate correct present checks for implicit maps of optional allocatables (#138210 ) Currently, we do not generate the appropriate checks to check if an optional allocatable argument is present before accessing relevant components of it, in particular when creating bounds, we must generate a presence check and we must make sure we do not generate/keep an load external to the presence check by utilising the raw address rather than the regular address of the info data structure. Similarly in cases for optional allocatables we must treat them like non-allocatable arguments and generate an intermediate allocation that we can have as a location in memory that we can access later in the lowering without causing segfaults when we perform "mapping" on it, even if the end result is an empty allocatable (basically, we shouldn't explode if someone tries to map a non-present optional, similar to C++ when mapping null data).	2025-05-09 13:57:45 +02:00
Joseph Huber	dbe070eb3e	[Offload] Fix PowerPC builds that pass -mcpu (#138327 ) Summary: Another hacky fix done until https://github.com/llvm/llvm-project/pull/136729 lands. This time for `-mcpu`.	2025-05-06 14:14:16 -05:00
Joseph Huber	dfcb8cb2a9	[OpenMP] Add pre sm_70 load hack back in (#138589 ) Summary: Different ordering modes aren't supported for an atomic load, so we just do an add of zero as the same thing. It's less efficient, but it works. Fixes https://github.com/llvm/llvm-project/issues/138560	2025-05-05 16:33:41 -05:00
Ye Luo	dcb43307ce	[Offload] Fix dependency issue #126143 in CMake	2025-05-05 00:38:48 -05:00
Michał Górny	d1e38eab95	[offload] Fix enabling unittests in standalone builds (#138418 ) Modify the unittest logic in offload to only look for `third-party/unittest` directory when `llvm_gtest` is not provided by LLVM itself (in-tree or installed). This makes it possible to run unittests in sparse checkouts without the `third-party/unittest` tree. While at it, also make sure `LLVM_THIRD_PARTY_DIR` is actually set while performing standalone builds. The logic is copied from `compiler-rt`. --------- Co-authored-by: Joseph Huber <huberjn@outlook.com>	2025-05-03 20:06:14 +02:00
Ross Brunton	f6ac5276ee	[Offload] Ensure all `llvm::Error`s are handled (#137339 ) `llvm::Error`s containing errors must be explicitly handled or an assert will be raised. With this change, `ol_impl_result_t` can accept and consume an `llvm::Error` for errors raised by PluginInterface that have multiple causes and other places now call `llvm::consumeError`. Note that there is currently no facility for PluginInterface to communicate exact error codes, but the constructor is designed in such a way that it can be easily added later. This MR is to convert a crash into an error code. A new test was added, however due to the aforementioned issue with error codes, it does not pass and instead is marked as a skip.	2025-05-02 07:37:19 -05:00
Joseph Huber	a60984ec8d	[Offload] Add 'Maintainers.md' file for offload (#138177 ) Summary: The offload project lacks a maintainers file. Adding it with myself and Johannes as the still active maintainers.	2025-05-01 14:06:33 -05:00
Ross Brunton	49941749a8	[offload] Don't print device path during configure (#138109 )	2025-05-01 07:44:26 -05:00
Callum Fare	7bc16a0f63	[Offload] Adding missing Offload unit tests for event entry points (#137315 ) A couple of liboffload entry points were missed out from the tests, and unsurprisingly a crash in one of them made it in. Add the tests and fix the unchecked error in `olDestroyEvent`.	2025-04-30 09:06:00 -05:00
Callum Fare	6022a5214b	[Offload] Add check-offload-unit for liboffload unittests (#137312 ) Adds a `check-offload-unit` target for running the liboffload unit test suite. This unit test binary runs the tests for every available device. This can optionally filtered to devices from a single platform, but the check target runs on everything. The target is not part of `check-offload` and does not get propagated to the top level build. I'm not sure if either of these things are desirable, but I'm happy to look into it if we want. Also remove the `offload/unittests/Plugins` test as it's dead code and doesn't build.	2025-04-29 11:21:59 -05:00

1 2 3 4 5 ...

311 Commits