Reverse engineering the storage format for an undocumented database
Posted by pintprint 4 days ago
Comments
Comment by elias1233 2 days ago
Comment by xpct 2 days ago
Comment by whsoul 1 day ago
during developing Kafka search tool, it have to support many kind of decode formats ( such as Avro + schemaRegistry, protobuf, ConnectJson(json with schema), and so on ). it recommend a type of deserializer by looking at the leading bytes. After user select final decision, it is ok if error occurs during first data deserialize, But many case it’s not occurred error, just pass through with plausibly wrong parsed data. So my solution was display 10 pre decoded sample to user, and make user select one deserializer by eye.
ConnectJson is reliable, because it has data and schema embedded in every message ( but large ). In my case using schema-registry, server1 has schema no 10, and server2 has schema no 10 but it’s different. I make a mistake select wrong server 2, data parsing is progressed but plausibly wrong data It took me a long time to find and correct some thing wrong.
Comment by gotosun1 1 day ago
Comment by tancop 1 day ago
Comment by pintprint 20 hours ago
Comment by quietraster 1 day ago
Comment by pintprint 20 hours ago
But of course it was a fun process and very satisfying when the normalizer actually started to work.
Comment by russelmelroy 14 hours ago