没有找到合适的产品?
联系客服协助选型:023-68661681
提供3000多款全球软件/控件产品
针对软件研发的各个阶段提供专业培训与技术咨询
根据客户需求提供定制化的软件开发服务
全球知名设计软件,显著提升设计质量
打造以经营为中心,实现生产过程透明化管理
帮助企业合理产能分配,提高资源利用率
快速打造数字化生产线,实现全流程追溯
生产过程精准追溯,满足企业合规要求
以六西格玛为理论基础,实现产品质量全数字化管理
通过大屏电子看板,实现车间透明化管理
对设备进行全生命周期管理,提高设备综合利用率
实现设备数据的实时采集与监控
利用数字化技术提升油气勘探的效率和成功率
钻井计划优化、实时监控和风险评估
提供业务洞察与决策支持实现数据驱动决策
原创|行业资讯|编辑:龚雪|2015-03-11 13:09:15.000|阅读 247 次
概述:很多人听说过企业搜索,但很少有人知道企业搜索的底层其实是文档过滤器。
# 界面/图表报表/文档/IDE等千款热门软控件火热销售中 >>
IF YOU LOOKED at a Microsoft Word file in binary format (as a search engine needs to review it), the file structure is so complex as to make it nearly impossible to pick out the text. In fact, MS Word documents include not only body text but also fields and often even hidden meta data. And MS Word files can have a nested structure, embedding multiple layers of other documents within the Word file.
Delving through these levels of complexity requires a programmatic implementation embedding a deep understanding of file structure. That is the job of document filters.
Document filters are a dynamic component. Every update, for example, that Microsoft makes to the MS Word format requires an adjustment to the document filters going forward, while still preserving backward compatibility with existing Word files.
One leading supplier of enterprise and developer text search software, dtSearch Corp., has spent over two decades building its own document filters. And the company continually upgrades its document filters to correspond with the release of new data formats.
In addition to Word, other MS Office file types that dtSearch supports include PowerPoint, Excel, Access, and OneNote. The document filters also support PDF, RTF, OpenOffice, HTML, XML, CSV, and many other file types, along with compression formats like RAR, ZIP, and GZIP/TAR. And the dtSearch document filters support recursively embedded versions of files, such as a Word file embedded in an Excel file contained in a ZIP attachment.
The dtSearch document filters can also support browser-compatible images in files, including recursively embedded files. The document filters further include Unicode support covering hundreds of international languages.
With so much data now in emails, the dtSearch document filters also support email formats like MS Outlook, Exchange, and Thunderbird. And support extends beyond the email body and meta data to cover multi-layered nested attachments, including recursively-embedded images.
The dtSearch Engine APIs can also work with database data like SQL. While SQL itself is not a file format, it can include BLOB data consisting of embedded documents. The same integrated support for recursively embedded documents, meta data, images, and the like apply to this BLOB data.
Finally, the dtSearch Spider supports static and dynamic Web data (SharePoint, PHP, ASP.NET, CMS, etc.). This data can further consist of (or simply embed) document data such as HTML, PDF, XSL/XML, or even Office files, all of which require the document filters.
dtSearch enterprise and developer products can index more than a terabyte of data in a single index. A single index can span multiple file directories, emails and attachments, online data, and other databases. The products can create and search any number of indexes.
After indexing, the product line supports highly concurrent, multithreaded searching. Indexed search time is typically less than a second, even across terabytes of data. dtSearch products offer more than 25 search options.
For federated searching, dtSearch products support integrated relevancy ranking across both online and offline repositories. Following a search, the document filters enable hit-highlighting of federated search content.
In the dtSearch Engine, API filters and objects provide an even wider range of advanced data classification options. SDKs include native 64-bit and 32-bit APIs for C++, Java, and .NET (through current versions).
本站文章除注明转载外,均为本站原创或翻译。欢迎任何形式的转载,但请务必注明出处、不得修改原文相关链接,如果存在内容上的异议请邮件反馈至chenjj@ke049m.cn
文章转载自:慧都控件网Tech Soft 3D的HOOPS Exchange与HOOPS Access,还是Spatial的3D InterOp,它们都体现了当前工程软件领域在数据互操作技术上的发展趋势—— 即以 高精度几何解析、跨平台开放架构与可持续兼容性 为核心,构建从设计、仿真到制造的数字数据链。
在现代复杂系统开发过程中,需求管理是确保项目成功的关键环节。Sparx Systems公司的Enterprise Architect作为一款先进的UML建模和设计工具,其需求管理模块通过完整的追溯机制,为项目提供了从需求收集到设计实现、测试验证的全生命周期可追溯性解决方案,有效保障了项目交付质量与规范符合度。
在企业应用、报表系统或财务工具的开发中,生成规范、专业的 PDF 文档是常见需求。与其在代码中硬编码布局,不如使用模板来提高开发效率。模板不仅能加快开发进程,还能确保品牌视觉与文档格式的一致性。本文将介绍如何使用 Spire.PDF for .NET 在 C# 中通过 HTML 模板 或 预设 PDF 模板 生成 PDF 文档,无论是需要动态布局还是快速替换占位符,都能灵活应对。
近日,全球知名的文档与图像处理组件Aspose正式推出 25.10 版本!本次更新覆盖 Words、Cells、PDF、Imaging、CAD、PSD、OCR 等多条产品线,重点聚焦性能提升、格式兼容性优化以及跨语言平台的统一支持,为开发者提供更高效、更稳定的企业级文档处理体验。
全球领先的文本检索工具,支持在千兆字节数量级的数据源中进行搜索。
dtSearch Network with Spider全球领先的文本检索工具,支持在千兆字节数量级的数据源中进行搜索。
dtSearch Web with Spider全球领先的文本检索工具,能够快速地将大量的搜索内容即时发布到基于IIS的Web站点上。
dtSearch Publish全球领先的文本检索工具,能够为CD/DVD publishing提供强大的功能。
dtSearch Engine超过20年的全球领先的文本检索控件,使开发者为应用程序快速添加文本查检索功能。
服务电话
重庆/ 023-68661681
华东/ 13452821722
华南/ 18100878085
华北/ 17347785263
客户支持
技术支持咨询服务
服务热线:400-700-1020
邮箱:sales@ke049m.cn
关注我们
地址 : 重庆市九龙坡区火炬大道69号6幢