WEBVTT 00:04.230 --> 00:09.840 So now you understand that the topic may have multiple partitions, those partitions are spread among 00:09.840 --> 00:17.310 different brokers and managers inside of every politician must have unique of that number and of set 00:17.310 --> 00:17.900 numbers. 00:17.910 --> 00:19.420 I started from zero. 00:19.440 --> 00:25.650 So our first message inside of every politician has of set number zero new messages appended at the 00:25.650 --> 00:28.880 end of existing messages inside of every partition. 00:29.310 --> 00:35.850 But there is a problem if any of the brokers fails is partition simply disappears and messages that 00:35.860 --> 00:38.310 were stored in that partition also disappear. 00:39.030 --> 00:43.230 That's why it is possible to create replicas of partitions. 00:43.500 --> 00:43.860 And in. 00:44.460 --> 00:45.460 Let me talk about that. 00:45.870 --> 00:51.600 So here you see a diagram with three brokers, brokers zero, broker one and broker two. 00:51.870 --> 00:56.850 And there is topic A and notice that in this topic there is partition zero. 00:56.860 --> 01:02.910 Let's suppose that there was just single partition so far and you see the same partition partition zero 01:02.910 --> 01:06.120 that was created on brokers zero and broker to. 01:07.300 --> 01:11.210 But those partitions are created as follows partitions. 01:11.500 --> 01:21.340 So this broker and this broker, followers and broker one is leader for this partition and job of followers 01:21.490 --> 01:27.430 is simply to get new messages from the leader and ride them into specific partition. 01:27.620 --> 01:28.390 That's all. 01:28.820 --> 01:33.640 They don't accept any riden requests for this partition from producers. 01:34.120 --> 01:36.580 Also, they don't serve consumers. 01:37.300 --> 01:45.250 Producers and consumers communicate only with Broker One when they want to write or read from Partition 01:45.250 --> 01:45.810 Zero. 01:46.240 --> 01:53.140 And that's why partition zero on this broker one is called Leader Partition, or this broker is called 01:53.140 --> 01:55.490 leader for this particular partition. 01:55.930 --> 02:02.950 And again, when, for example, new message arrives to partition zero and broker one, it accepts write 02:02.950 --> 02:11.080 request from producer creates new message here in its partition zero and replicates this same message. 02:11.090 --> 02:18.130 Message number two, for example, to partition zero and broker a zero here and to broker to partition 02:18.130 --> 02:19.140 zero here. 02:19.510 --> 02:24.820 And same message number two is created here, here and here. 02:25.270 --> 02:29.050 And now there are two replicas of the same message. 02:29.620 --> 02:38.710 And if broker one fails now one of those brokers is our broker or broker to become leader for the same 02:38.710 --> 02:45.370 partition, partition zero and no, for example, Broker Zero will still continue the operation and 02:45.370 --> 02:51.740 will accept new writing requests to this partition and accept Reid's request from consumers. 02:52.480 --> 02:56.770 That's how you are able to achieve for tolerance using replication. 02:57.040 --> 03:03.130 Again, Main India behind this application is that leader performs most operations. 03:03.400 --> 03:11.680 It communicates with producers and consumers and also it sends copies of every message to all followers. 03:12.490 --> 03:16.030 Followers simply sit and relax and wait for new messages. 03:16.090 --> 03:21.990 And as soon as new message arrives from the leader, their job is to write this message to partition. 03:22.000 --> 03:23.490 And that's all, nothing else. 03:24.250 --> 03:31.510 That's why you need to blend resources of the brokers accordingly, because if there will be multiple 03:31.510 --> 03:39.090 replicas for every partition, for every topic then logged on brokers who will be responsible for lead, 03:39.160 --> 03:41.960 the role will increase hugely increase. 03:42.100 --> 03:43.840 So please keep all of that in mind. 03:44.440 --> 03:49.000 All that is recommended to create the at least the two replicas. 03:49.210 --> 03:56.170 I mean, if you want to create the full all around architecture, you need to have at least multiple 03:56.170 --> 03:56.710 brokers. 03:56.710 --> 04:03.160 And if you want to create two copies of every message, you need to have at least three brokers and 04:03.160 --> 04:07.660 you need to configure a so-called replication factor on a topic level. 04:07.690 --> 04:13.720 Basically, it means that you are not able to set up a replication factor for every specific partition 04:13.720 --> 04:14.590 inside of the topic. 04:14.860 --> 04:21.040 Instead, you must configure it on a topic level basis and by default, replication factor is a set 04:21.040 --> 04:21.640 to one. 04:21.820 --> 04:27.070 And that means that every message is stored only once, only on one broker. 04:27.550 --> 04:33.460 And if you want to achieve such architecture where every message from a leader will be replicated twice 04:33.670 --> 04:38.170 to two different brokers, then replication a factor must be set to three. 04:38.530 --> 04:43.930 And that means that every message will be stored three times on three different brokers. 04:44.830 --> 04:50.380 It is actually a recommended number for production environments and you should not go further. 04:50.380 --> 04:53.880 And the set of application de facto to a four or five and so on. 04:54.010 --> 05:00.490 So three is completely enough and this number will tolerate shutdown down of two brokers of three that 05:00.610 --> 05:08.090 we're storing replicas of every message in this case, broker one and broker two may fail and broker 05:08.090 --> 05:13.630 a zero will still continue operation and it will still keep old messages in partition zero. 05:14.290 --> 05:20.890 And if I go back to this diagram where we had three partitions, partition zero, one and two, and 05:20.890 --> 05:28.030 if you stop replication of fact on three, then basically we will have nine partitions instead of three, 05:28.390 --> 05:34.690 but six of them will be, let's say, passive, and they will be created as four in partitions and the 05:34.840 --> 05:39.700 leaders will write messages, replicate messages to those following partitions. 05:39.910 --> 05:47.320 But as you see, with replication factor in place, quantity of partitions basically multiplies by replication 05:47.320 --> 05:47.900 of factor. 05:48.010 --> 05:51.610 So here now we have three partitions, but with a replication factor. 05:51.610 --> 05:53.860 Three, we need to have nine partitions. 05:54.560 --> 06:00.570 OK, that's all about Leedham, leader of partition, and that's all about the replication. 06:01.000 --> 06:01.990 I hope it is clear. 06:01.990 --> 06:06.140 And if you use multiple brokers, multiple partitions in each topic and. 06:06.710 --> 06:13.610 A factor you might achieve really nice results and you may build real nice for tolerant zylon to Kafka. 06:13.620 --> 06:16.630 Kloster again, this all about this topic. 06:16.640 --> 06:19.010 And next, let's going to talk about control. 06:19.220 --> 06:22.510 I'll explain you what is controller and what is his job. 06:22.610 --> 06:23.390 So we'll see you next.