如何使用pdfbox从pdf中删除可选内容组及其内容? [英] How to delete an optional content group alongwith its content from pdf using pdfbox?
本文介绍了如何使用pdfbox从pdf中删除可选内容组及其内容?的处理方法,对大家解决问题具有一定的参考价值,需要的朋友们下面随着小编来一起学习吧!
问题描述
我已经实现了从pdf删除图层的功能,但是问题是,我在图层上绘制的内容没有被删除.这是我用来删除图层的代码:>
I have implemented functionality to delete the layer from pdf, but the problem is that, the content that I had drawn on the layer, does not get delete.Here is the code that I am using to delete the layer:
PDDocumentCatalog documentCatalog = doc.getDocumentCatalog();
PDOptionalContentProperties ocgProps = documentCatalog.getOCProperties();
PDOptionalContentGroup ocg = ocgProps.getGroup(markupLayerName);
COSDictionary ocgsDict = (COSDictionary)ocgProps.getCOSObject();
COSArray ocgs = (COSArray)ocgsDict.getItem(COSName.OCGS);
int indexToBeDeleted = -1;
for (int index = 0; index < ocgs.size(); index++)
{
COSBase o = ocgs.get(index);
COSDictionary ocgDict = ToCOSDictionary(o);
if (ocgDict.getString(COSName.NAME) == markupLayerName)
{
indexToBeDeleted = index;
break;
}
}
if (indexToBeDeleted >= 0)
{
cgs.remove(indexToBeDeleted);
ocgsDict.setItem(COSName.OCGS, ocgs);
documentCatalog.setOCProperties(new PDOptionalContentProperties(ocgsDict));
}
推荐答案
要删除标记数据,我不得不修改PDPage的内容.我只是在内容中搜索了BDC和EMC对,然后搜索该对是否属于有关的层,如果是这样,那么我将从内容中删除该部分.下面是我使用的C#代码:
To delete the markup data , I hade to modify the PDPage's content.I just searched the contents for BDC and EMC pair, and then searched whether that pair belongs to the concerned layer, if so then I delete that part from the contents.Below is the C# code that I used:
PDPage page = (PDPage)doc.getDocumentCatalog().getPages().get(pageNum);
PDResources resources = page.getResources();
PDFStreamParser parser = new PDFStreamParser(page);
parser.parse();
java.util.Collection tokens = parser.getTokens();
java.util.List newTokens = new java.util.ArrayList();
List<Tuple<int, int>> deletionIndexList = new List<Tuple<int, int>>();
object[] tokensArray = tokens.toArray();
for (int index = 0; index < tokensArray.Count(); index++)
{
object obj = tokensArray[index];
if (obj is COSName && (((COSName)obj) == COSName.OC))
{
int startIndex = index;
index++;
if (index < tokensArray.Count())
{
obj = tokensArray[index];
if (obj is COSName)
{
PDPropertyList prop = resources.getProperties((COSName)obj);//Check if the COSName found is the resource name of layer which contains the markup to be deleted.
if (prop != null && (prop is PDOptionalContentGroup))
{
if (((PDOptionalContentGroup)prop).getName() == markupLayerName)
{
index++;
if (index < tokensArray.Count())
{
obj = tokensArray[index];
if (obj is Operator && ((Operator)obj).getName() == "BDC")//Check if the token specifies the start of markup
{
int endIndex = -1;
index++;
while (index < tokensArray.Count())
{
obj = tokensArray[index];
if (obj is Operator && ((Operator)obj).getName() == "EMC")//Check if the token specifies the end of markup
{
endIndex = index;
break;
}
index++;
}
if (endIndex >= 0)
{
deletionIndexList.Add(new Tuple<int, int>(startIndex, endIndex));
}
}
}
}
}
}
}
}
}
int tokensListIndex = 0;
for (int index = 0; index < deletionIndexList.Count(); index++)
{
Tuple<int, int> indexes = deletionIndexList.ElementAt(index);
while (tokensListIndex < indexes.Item1)
{
newTokens.add(tokensArray[tokensListIndex]);
tokensListIndex++;
}
tokensListIndex = indexes.Item2 + 1;
}
while (tokensListIndex < tokensArray.Count())
{
newTokens.add(tokensArray[tokensListIndex]);
tokensListIndex++;
}
PDStream newContents = new PDStream(doc);
OutputStream output = newContents.createOutputStream(COSName.FLATE_DECODE);
ContentStreamWriter writer = new ContentStreamWriter(output);
writer.writeTokens(newTokens);
output.close();
page.setContents(newContents);
这篇关于如何使用pdfbox从pdf中删除可选内容组及其内容?的文章就介绍到这了,希望我们推荐的答案对大家有所帮助,也希望大家多多支持IT屋!
查看全文