A command logged by the exec package could not be copied out of the log
and re-run. Creating the RGW admin ops user, for example, logged an
unquoted "--display-name RGW Admin Ops User" and an unquoted "--caps
accounts=*;buckets=*;users=*;usage=read;metadata=read;zone=read", which
a shell reads as four arguments and six commands respectively. Render
these lines with FormatCommand instead.
cmd-reporter also named the command twice, because exec.Cmd already
carries argv[0] in Args and the log line prefixed Path to the whole of
Args.
The startup line drops the quotes that wrapped the entire argument list,
now that each argument carries its own; both in-tree samples of that
line are updated to match.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
If osd-replacement fails during provisioning of the new device by the
osd-prepare job (can be the case if the new device is faulty), then
ceph-volume will run its internal rollback logic, which purges the
reserved osd id and its CRUSH position. After the faulty device is
replaced with a working one, rook will reprovision it correctly and
reuse the metadata device slot, but as a new OSD, so data will likely
be rebalanced. Ceph hands out the lowest free id, so the new OSD often
takes the same number back, and then the create path recreates the
deployment by itself. Otherwise, and until a working device arrives,
the downscaled deployment of the destroyed osd is left behind. This
commit handles that case by deleting downscaled ready-for-swap osd
deployments without a corresponding OSD in the osd dump.
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
osd-replacement was relying on exising ceph.rook.io/do-not-reconcile
label for fencing osd destroy process owned by osd health goroutine from
controller. However, this label can be used but other components like
rook krew maintenance plugin and cannot be owned by osd-replacement
process. Added a separate osd.rook.io/replace-in-progress annotation for
that purpose.
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
After some markdown link fixes,
make docs-build failed due to some non-existing anchors in documentation
sources.
This changes fixes the docs-build by rectifying anchors.
Assisted-by: IBM-Bob
Signed-off-by: Michael Adam <obnox@samba.org>
This change fixes some of the broken links in the markdown
documentatoion sources found by the markdown link checker.
Assisted-by: IBM Bob
Signed-off-by: Michael Adam <obnox@samba.org>
updated the external cluster doc, and helm and manifest
install to make use of csi operator as it is the default
offering now
Signed-off-by: parth-gr <partharora1010@gmail.com>
With the default version now being v20.2.1, also update
all of the examples, documentation, and default CI to run
with that version
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
when cluster is deployed in namespace other than rook-ceph,
csi-operator.yaml should also follow similar approach on how
rook operator.yaml and other yaml are updated to deploy in
alternate namespace using the comment `# namesapce: operator`.
Signed-off-by: subhamkrai <srai@redhat.com>
Implement CephX key rotation for Rook's client.admin user.
Admin user rotation is risky, so this has been tested extensively both
in unit tests as well as by manually injecting failures during runtime.
In testing, all failures were able to be recovered by the recovery
routine.
A mutex is also added to help ensure that two simultaneous admin key
rotation processes cannot be running simultaneously for any given
namespace. The mutex is tested in unit tests, and it was verified during
runtime via manual testing.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
Adds .spec.cephConfigFromSecret to CephCluster for loading
Ceph config parameters from a Kubernetes Secret.
Signed-off-by: Patryk Rostkowski <patrostkowski@gmail.com>
The change creates and manages an Endpoints resource
with the current Ceph monitor (mon) IPs, allowing
clients to resolve mon IPs via DNS without relying
on the rook-ceph-mon-endpoints ConfigMap.
Signed-off-by: Patryk Rostkowski <patrostkowski@gmail.com>
Implements #14733. Allows to set IDs of external mons to
Cluster CRD. Rook will not remove external mons from quorum
and will add external mon addresses to mon endpoints.
Use-case for external mon is to maintain quorum for 2-AZ
k8s cluster in case of zone outage.
Signed-off-by: Artem Torubarov <artem.torubarov@clyso.com>
The cephConfig settings in the CephCluster CR have been
stable and there are no planned changes, so remove the
experimental documentation indicator. Also, clarify
the usage and the precedence of the ceph config
options.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
mkdocs uses a markdown renderer that is hardcoded to 4 spaces per tab
for detecting indentation levels, including ordered- and
unordered-lists. Since we cannot easily change the renderer, begin using
a markdown linter in CI that will fail if official docs do not adhere to
the spacing rules.
As a starting point, the markdownlint config does not begin with the
default set of checks, which might overwhelm attempts to fix them.
Instead, focus on list-tab-spacing rules and a few other highly useful
checks.
markdownlint also has some gaps in its abilities that allow common Rook
doc issues to pass acceptance. However, it allows creating custom
linting plugins. Create 2 such linting plugins to check 2 things:
- all doc lines (except code blocks) must be aligned to a 4-space
boundary, without exception. This ensures that markdown will render
correctly with mkdocs. This unfortunately makes it possible to create
lists that are internally aligned strangely.
- admonitions must all follow the same format of
```
!!! header
body
```
For the strange lists, this is allowed and renders correctly, but it
looks strange:
```md
- first bullet
- second bullet
still second bullet
- third bullet
has a paragraph
of text inside
- last bullet
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
Users on Microsoft Azure can make use of the Azure key
vault service rather than replying on any third party
service for KMS.
Signed-off-by: sp98 <sapillai@redhat.com>
For now we are using the operator namespace name
as the prefix for the csi driver, This PR provides
an option for the users if someone wants to have
their own prefix for the csi driver, if someone tries
to change the prefix for existing csi driver rook
operator will fail to reconcile the csi driver.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
This implements the "Ceph Config via Ceph Cluster CRD" design document
as a `cephConfig:` structure on the CRD.
This also fixes the `yq` commands used to manipulate the
`cluster-test.yaml` that caused CI issues for this PR and potentially
unknowingly others.
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
Let's use the same technique based on 'sed' while employing two
different namespaces for the cluster and the operator when creating
the second cluster.
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
When a user attempts to adhere to the current documentation to use alternative
namespaces for both the operator and the cluster, they fail because
our common YAML file only has a single namespace for the cluster. This
change adds specific instructions to create operator namespace.
Closes: https://github.com/rook/rook/issues/13079
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
Change how Rook detects network CIDRs for Multus networks. The IPAM
configuration is only defined as an arbitrary string JSON blob with a
"type" field and nothing more. Rook's detection of CIDRs for whereabouts
had already grown out of date since the initial implementation.
Additionally, Rook did not support DHCP IPAM, which is a reasonable
choice for users. And more, Rook did not support CNI plugin chaining,
which further complicates NADs. Based on the CNI spec, network chaning
can result in any changes to network CIDRs from the first-given plugin.
All these problems make it more and more difficult for Rook to support
Multus by inspecting the NAD itself to predict network CIDRs. Instead,
it is better for Rook to treat the CNI process as a black box. To
preserve legacy functionality of auto-detecting networks and to make
that as robust as possible, change to a canary-style architecture like
that used for Ceph mons, from which Rook will detect the network CIDRs
if possible.
Also allow users to specify overrides for CIDR ranges. This allows Rook
to still support esoteric and unexpected NAD or network configurations
where a CIDR range is not detectable or where the range detected would
be incomplete. Because it may be impossible for Rook to understand the
network CIDRs wholistically while residing only on a portion of the
network, this feature should have been present from Multus's inception.
Improving CIDR auto-detection and allowing users to specify overrides
for auto-detected CIDRs rounds out Rook's Multus support for CephCluster
(core/RADOS) installations. No further architectural changes should be
needed for CephClusters as regards application of public/cluster network
CIDRs for Multus networks.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
updated troubleshoot doc using common issues to reference
krew plugin and added krew section for osd pugre.
Signed-off-by: subhamkrai <srai@redhat.com>
- adds description on how to configure default pg_num and pgp_num
parameters on per pool basis while deploying rook based ceph cluster.
- advices user to declare them during initial startup.
uptil now, docs only covered cli way of configuring them.
Closes: https://github.com/rook/rook/issues/10361
Signed-off-by: Deepika Upadhyay <deepika@koor.tech>
With octopus coming to end of life, we remove support from
Rook for deploying Ceph Octopus and assume a min version of
Pacific v16. Any checks for octopus or earlier are removed
from the reconciles since they are obsolete.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The cluster CR doc is extremely long and in need of refactoring.
Now we have four new sub-topics for a host-based cluster,
PVC-based cluster, stretch cluster, and an external cluster.
Each of those topics has its own description and examples.
The main cluster CR topic still contains the details about
all the possible cluster CR settings.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>