Until now, an object store would create all the necessary
metadata pools and the data pool that were exclusively
for its own object store. When isolation between object
stores is necessary, this would cause many pools and
PGs to be created in the cluster, which was not
manageable.
Now one set of pools can be created to be shared
by any number of object stores. The metadata and data
between each object store is isolated by
RADOS namespaces, which by design will keep the
data safe for multi-tenancy.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
if Multus is enabled the clusterinfo should be updated with
network as multus as to run the ceph cmds in remote
executor
Signed-off-by: parth-gr <paarora@redhat.com>
User can define his desired endpoint list in Zone CR so that it will
overwrite the default service name for rgw.
Resolves#6432
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
Signed-off-by: Jiffin Tony Thottan <jthottan@redhat.com>
The CSI package needs to load clusterInfo, today this code is in the mon
package which makes the call of LoadClusterInfo impossible without
having a circular import.
Signed-off-by: Sébastien Han <seb@redhat.com>
adding observedGeneration field in the cephcluster cr
status for having better control on reconciling,
as observedGeneration field will be updated by the controller
Closes: https://github.com/rook/rook/issues/9673
Signed-off-by: parth-gr <paarora@redhat.com>
The cluster CR can be forcefully deleted and cleanup the
cluster resources if the yes-really-destroy-data policy
is set on the CR. In this case, the other controllers should
treat the cluster CR as not existing and allow the finalizers
to be removed on those resources if they are requested for
deletion.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit adds context parameter to utilities in opcontroller to
remove context.TODO use in opcontroller. By this, we can handle
cancellation of reconcilers in a fine-grained way.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
Sometimes `radosgw-admin` succeeds after showing logs to stderr. We should
skip non-json strings if parsing output as json.
Here is an example.
```
2021-02-26 04:10:44.190418 I | op-bucket-prov: creating Ceph user "ceph-user-aSzNqgE7"
E0226 04:11:37.901310 8 controller.go:199] error syncing 'logging/loki-bucket': error provisioning bucket: Provision: can't create ceph user: error creating ceph user "ceph-user-aSzNqgE7": failed to unmarshal json. 2021-02-26T04:11:21.425+0000 7f6714be4980 1 robust_notify: If at first you don't succeed: (110) Connection timed out
2021-02-26T04:11:21.426+0000 7f6714be4980 0 ERROR: failed to distribute cache for ceph-hdd-object-store.rgw.meta:users.uid:ceph-user-aSzNqgE7
2021-02-26T04:11:32.168+0000 7f6714be4980 1 robust_notify: If at first you don't succeed: (110) Connection timed out
2021-02-26T04:11:32.168+0000 7f6714be4980 0 ERROR: failed to distribute cache for ceph-hdd-object-store.rgw.meta:users.keys:23Z8GUEXR0TJDO86PSJR
{
"user_id": "ceph-user-aSzNqgE7",
"display_name": "ceph-user-aSzNqgE7",
...
"mfa_ids": []
}: invalid character '-' after top-level value: failed to unmarshal json.
```
In this case, some logs like "robust_notify:..." was shown in stderr.
Unmarsharing was failed due to tried to parse these logs as json.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
Found by running the following command:
codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H
Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
Pools in stretch clusters will all use the same crush rule that
is generated by the operator when configuring the cluster for
stretch mode, with replica 4 and two replicas per failure domain.
Pools cannot create new crush rules, so we require that replica be
4 when the pools is created, EC is not allowed, and no new
crush rule will be created.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The CRUSH map does not allow duplicate values (nor keys) for the
labels. That means that if the currently hardcoded label value
`default` occurs anywhere else in the tree (for example, because
some of the topology labels provided by the cloud provider use it),
the cluster cannot sucessfully spawn.
By giving the user control over the label value used for the root
CRUSH map label, it is possible to adapt the cluster to this
situation.
To be able to actually use the `default` value, the root=default
node which is created by default by ceph itself has to be removed;
since that node is accompanied by a replicated crush rule, we have
to remove that rule, too.
The integration tests are modified to sometimes use the custom
crushRoot (based on an arbitrary criterium) to get coverage across
the various suites. Separate tests could be added, but were not
deemed necessary at this point.
Fixes#4993.
Signed-off-by: Jonas Schäfer <jonas.schaefer@cloudandheat.com>
Add unit tests for when the "Zone" is configured
in an object store so that the object store joins
the CephObjectZone and the zone's corresponding
multisite configuration.
Signed-off-by: Ali Maredia <amaredia@redhat.com>
this commit will enable one more linter ineffassign
in golangci-lint.
This linter throws an error when variable is assigned and never used.
`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
- Add realm/zone group/zone to Object Context,
so that any call to runAdminCommand has the
same realm, zone group, and zone as the object
store that the call is going to.
- Also modify the delete code for the object-store
using multisite to remove the endpoints of an
object store from a zone instead of the normal
clean-up
- Added more debug logging all around the object-store
code related to multisite
- files generated by rerun of `make codegen`
- change back the edit on the Copyright in object.go
Signed-off-by: Ali Maredia <amaredia@redhat.com>
When zones are created the ceph RGWs inside the
those zones should be using pools with the zones
name not the object-store's name
Signed-off-by: Ali Maredia <amaredia@redhat.com>
This commit:
- adds the pull section on the CephObjectRealm spec
to enables realms to be pulled instead of created.
- creates the system user and generates the access
key and secret key for the user and zones in a realm.
- adds "omitempty" to fields in multisite CRs that
are not required
Signed-off-by: Ali Maredia <amaredia@redhat.com>
Name for requeue time variables were the same for
the object-zone and the object-zonegroup.
This commit makes them more descriptive and
accurate
Signed-off-by: Ali Maredia <amaredia@redhat.com>
The masterArg flag cannot be "" so a different
command without masterArg must be run for zone
and zone group creation
Signed-off-by: Ali Maredia <amaredia@redhat.com>
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.
Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit contains CR implementations for:
CephObjectRealm
CephObjectZoneGroup
CephObjectZone
Also there are changes made to the objectstore
to add rgws in the object-store to zones and
zone groups in a multisite configuration and
the removal of the --default parameter for any
realms/zonegroups/zones that are created.
Signed-off-by: Ali Maredia <amaredia@redhat.com>