Telugu newspapers lo nunchi text ni extract cheyyali ani anukuntunna. Meer evar aina ilantivi chesara?

Hello andhariki,

Nenu telugu e-newspapers(like eenadu, andhrajyothi, etc) nunchi only news text ni extract cheyyali ani anukuntunna.

So main problem ekkada ante layout detection, article grouping and reading order.

Oka headline ki inko article lo unna content vasthundhi.

So 3 days nunchi try chesthunna, response sarigga vasthaledhu. Meeru em anna suggestions ivvagalara?

Plzzz 😭😭

reddit.com
u/gnanatejadiviti_22 — 1 day ago
▲ 1 r/telugu

Telugu newspapers lo nunchi text ni extract cheyyali ani anukuntunna. Meer evar aina ilantivi chesara?

Hello andhariki,

Nenu telugu e-newspapers(like eenadu, andhrajyothi, etc) nunchi only news text ni extract cheyyali ani anukuntunna.

So main problem ekkada ante layout detection, article grouping and reading order.

Oka headline ki inko article lo unna content vasthundhi.

So 3 days nunchi try chesthunna, response sarigga vasthaledhu. Meeru em anna suggestions ivvagalara?

Plzzz 😭😭

reddit.com
u/gnanatejadiviti_22 — 1 day ago
▲ 3 r/Archivists+2 crossposts

What are the different ways to extract text from Telugu language Newspapers?

Hi all, I want to extract telugu text from e-newspapers. They are in pdf format and I'm converting each page into an image.

Passing PP-DocLayoutV2 on that for getting layouts. Croping those layouts and applying Tesseract on each crop. The problem is getting reading order. Basically I'm trying to get that reading order and it is not working well.

Any tips to get output like:

Page1:

Article1:

Headline:

Sub-headline:

Body:

Article2:

No ads, no images only text.

Also the connecting the continuety, like only short info is given in Page1 and other matter will be written in other papers. Example: (To be continued in page 6)

Any suggestions on these....🫠

reddit.com
u/gnanatejadiviti_22 — 6 days ago