Mar 21, 2024

The Implementation Principle Of Video Encoding

Leave a message

Video coding technology prioritizes the elimination of spatial and temporal redundancy. Next, let me introduce to you what method is used: the video is formed by continuous playback of different frames.

 

These frames are mainly divided into three categories: I frames, B frames, and P frames. I frame is an independent frame with all its own information, which is the most complete picture (occupying the largest space) and can be decoded independently without referring to other images. The first frame in the video sequence is always the I frame. P frame, "inter frame prediction encoding frame," requires reference to different parts of the previous I frame and/or P frame in order to encode. The P frame has a dependency on the previous P and I reference frames. However, the P-frame compression rate is relatively high and occupies less space.


B-frame, "Bidirectional Predictive Encoding Frame", with frames before and after being used as reference frames. Not only referring to the front, but also referring to the frames behind, so its compression rate is the highest, reaching 200:1. However, it is not suitable for real-time transmission (such as video conferencing) because it relies on subsequent frames.


By classifying frames, the size of the video can be significantly compressed. After all, the number of objects to be processed has significantly decreased (from the entire image to a region within the image).


If you grab a packet from the video stream, you can also see the information of the I frame
If we always calculate based on pixels, the amount of data will be relatively large. Therefore, we usually divide the image into different "blocks" or "MacroBlocks" and calculate them. A macro block is generally 16 pixels by 16 pixels.


The processing of I frame adopts intra frame encoding method, utilizing only the spatial correlation within the image of this frame. The processing of P-frames adopts inter frame encoding (forward motion estimation) while utilizing spatial and temporal correlations. Simply put, using motion compensation algorithms to remove redundant information

 

Send Inquiry