core/oga - oga

Commit Graph

Author	SHA1	Message	Date
Yorick Peterse	4346cc6ec3	Added a KAF fixture file This is an output document from one of the components of the OpeNER (http://www.opener-project.eu/) project.	2015-08-18 22:32:57 +02:00
Dan Fockler	be7bc8f423	make string comparison faster	2015-08-15 06:36:27 -07:00
Dan Fockler	fc38b39aa3	make string comparison a bit faster	2015-08-15 06:34:09 -07:00
Daniel Fockler	496811a23f	Fixes #127	2015-08-14 16:15:49 -07:00
Yorick Peterse	0977be81bb	Release 1.2.2	2015-08-14 14:49:26 +02:00
Yorick Peterse	9b98e75115	Use pop/push vs shift/unshift in each_node While the performance difference between the old and new approach is pretty much negligible, it's simply not needed to use #shift/#unshift here. Thanks to Mon_Ouie from the #ruby IRC channel for suggesting this.	2015-08-07 14:08:54 +02:00
Yorick Peterse	32b75bf62c	Reset the owner ivar in LRU#synchronize Without this the following could happen: 1. Thread A acquires the lock and sets the ownership to A. 2. Thread A yields and returns 3. Thread B tries to acquire the lock 4. At this exact moment Thread A calls the "synchronize" method again and sees that the "owner" variable is still set to Thread A 5. Both thread A and B can now access the underlying data in parallel, possibly leading to corrupted objects This can be demonstrated using the following script: require 'oga' lru = Oga::LRU.new(64) threads = 50.times.map do Thread.new do loop do number = rand(100) lru[number] = number end end end threads.each(&:join) Run this for a while on either JRuby or Rubinius and you'll end up with errors such as "ConcurrencyError: Detected invalid array contents due to unsynchronized modifications with concurrent users" on JRuby or "ArgumentError: negative array size" on Rubinius. Resetting the owner variable ensures the above can never happen. Thanks to @chrisseaton for bringing this up earlier today.	2015-08-03 21:47:40 +02:00
Jakub Pawlowicz	ed3cbe7975	Fixes #129 - lexing superfluous end tags. Prevents a superfluous end tag of a self-closing HTML tag from closing its parent element prematurely, for example: ```html <object><param></param><param></param></object> ``` (note <param> is self closing) being turned into: ```html <object><param/></object><param/> ```	2015-07-23 13:16:11 +01:00
Yorick Peterse	d69be182bd	Switch to Travis' container infrastructure Fixes #128	2015-07-22 17:57:09 +02:00
Yorick Peterse	5d0e8c99af	Release 1.2.1	2015-07-01 00:24:07 +02:00
Jakub Pawlowicz	6786dde3f0	Splits entity decoding into two steps. Doing decimal & hex decoding separately results in a nicer code which does not refer to string matching and slicing.	2015-06-30 17:56:26 +02:00
Jakub Pawlowicz	6fc3ef425b	Fixes #118 - decoding invalid entities. Previous regular expression was too greedy in terms of matching letters from outside of A-F hex scope, and matching letters when not in hex mode.	2015-06-30 17:56:26 +02:00
Yorick Peterse	8990a62224	Release 1.2.0	2015-06-30 10:58:05 +02:00
Yorick Peterse	565e3da176	Added encoding comment in elements_spec.rb This ensures that older Ruby versions don't poop their pants when running these specs.	2015-06-29 21:09:33 +02:00
Yorick Peterse	dde644cd79	Support for Unicode XML/HTML identifiers Technically HTML only allows for ASCII names but restricting that actually requires more work than just allowing it.	2015-06-29 21:08:01 +02:00
Laurence Lee	139985612b	Lexer test for elements with inline dots.	2015-06-29 20:55:48 +02:00
Laurence Lee	b7771ed5fe	Patch to Oga Upstream to fix "Oga Parser doesn't handle dots in XML Tags". Fixes rubyjedi/soap4r#7	2015-06-29 20:55:48 +02:00
Yorick Peterse	71960fff87	Added CSS :nth() pseudo class This is a Nokogiri extension (as far as I'm aware) but it's useful enough to also include in Oga. Selectors such as "foo:nth(2)" are simply compiled to XPath "descendant::foo[position() = 2]". Fixes #123	2015-06-29 20:51:38 +02:00
Yorick Peterse	6c8f9704b3	Prepare CHANGELOG for the upcoming 1.1.1 release	2015-06-29 19:01:18 +02:00
Yorick Peterse	d26b48feb4	CSS parsing support for commas The lexer already had the basic plumbing in place but apparently I completely forgot to also implement the required bits in the parser. Fixes #121	2015-06-29 18:59:01 +02:00
Yorick Peterse	aa1a6ca308	Fixed CHANGELOG date for 1.1.0	2015-06-29 16:52:27 +02:00
Yorick Peterse	dccf7a7e06	Release 1.1.0	2015-06-29 16:45:21 +02:00
Yorick Peterse	3b633ff41c	Relax support for HTML unquoted attribute values This allows for parsing of HTML such as: <a href=lol("javascript")></a> Here the "href" attribute would have its value set to: lol("javascript") Fixes #119	2015-06-29 16:35:48 +02:00
Yorick Peterse	daec6d151c	Corrected "explici" type Fixes #120	2015-06-29 15:14:04 +02:00
Yorick Peterse	fcd5573422	Updated changelog for Node#replace	2015-06-17 21:48:40 +02:00
Tero Tasanen	0b4791b277	Ability to replace a node with another node or string ``` element = Oga::XML::Element.new(:name => 'div') some_node.replace(element) ``` You can also pass a `String` to `replace` and it will be replaced with a `Oga::XML::Text` node ``` some_node.replace('this will replace the current node with a text node') ``` closes #115	2015-06-17 21:27:50 +03:00
Yorick Peterse	fe7e2e4f74	Use ternary "if" when encoding attribute values	2015-06-16 22:48:42 +02:00
Yorick Peterse	074b53c18c	Fix entity encoding of attribute values This ensures that single and double quotes are also encoded, previously they would be left as is. Fixes #113	2015-06-16 22:47:10 +02:00
Yorick Peterse	e0837fd44e	Fixed YARD parameter name	2015-06-16 11:14:22 +02:00
Yorick Peterse	b701a9bdd4	Release 1.0.3	2015-06-16 11:13:06 +02:00
Yorick Peterse	b6d34a406d	Removed redundant returns	2015-06-16 00:51:51 +02:00
Yorick Peterse	bcffd86c50	Removed usage of YARD @!attribute macros	2015-06-16 00:04:10 +02:00
Yorick Peterse	2c18a51ba9	Support for strict parsing of XML documents Currently this only disabled the automatic insertion of closing tags, in the future this may also disable other features if deemed worth the effort. Fixes #107	2015-06-15 23:53:11 +02:00
Yorick Peterse	4031c4f843	Nuked Oga::XML::Lexer#html This method was rather pointless since there's already a "html?" method.	2015-06-15 23:45:20 +02:00
Yorick Peterse	020cd34083	Mark the XML lexer class as private I was supposed to mark this as private for 1.0 but completely forgot.	2015-06-15 23:18:11 +02:00
Yorick Peterse	fd307a0fcc	Support HTML attributes without starting quotes This allows the lexer to process input such as: <a href=foo"></a> For XML input the lexer still expects properly opened/closed attribute values. Fixes #109	2015-06-08 06:46:08 +02:00
Yorick Peterse	a76286b973	Support for spaces around attribute equal signs This also takes care of making sure line numbers are incremented properly. Fixes #112	2015-06-08 06:34:49 +02:00
Yorick Peterse	6f115bd8e0	Contributing section on submitting changes It pains me that I have to write this in the first place (I was hoping people would figure this out). Sadly history has shown it's required to document this properly.	2015-06-07 20:13:16 +02:00
Yorick Peterse	af7f2674af	Decoding of entities with numbers This ensures that entities such as "½" are decoded properly. Previously this would be ignored as the regular expression used for this only matched [a-zA-Z]. This was adapted from PR #111.	2015-06-07 17:42:24 +02:00
Yorick Peterse	f814ec7170	Expanded CONTRIBUTING guide Now includes more notes about specs, line wrapping, writing Git commit messages and more.	2015-06-07 17:22:44 +02:00
Yorick Peterse	5951a6f187	Release 1.0.2	2015-06-03 06:53:31 +02:00
Yorick Peterse	4bfeea2590	Use require vs require_relative See ruby-ll commit b27fe7cc109a39184ac984405a1e452868f3fac9 for a more in-depth explanation of this.	2015-06-03 06:42:30 +02:00
Yorick Peterse	d2523a1082	Support whitespace in element closing tags Fixes #108	2015-05-25 13:41:17 +02:00
Yorick Peterse	d0d597e2d9	Allow script/template in various table elements Fixes #105	2015-05-23 10:46:49 +02:00
Yorick Peterse	5182d0c488	Correct closing of unclosed, nested HTML elements Previous HTML such as this would be lexed incorrectly: <div> <ul> <li>foo </ul> inside div </div> outside div The lexer would see this as the following instead: <div> <ul> <li>foo</li> inside div </ul> outside div </div> This commit exposes the name of the closing tag to XML::Lexer#on_element_end (omitted for self closing tags). This can be used to automatically close nested tags that were left open, ensuring the above HTML is lexer correctly. The new setup ignores namespace prefixes as these are not used in HTML, XML in turn won't even run the code to begin with since it doesn't allow one to leave out closing tags.	2015-05-23 09:59:50 +02:00
Yorick Peterse	8172de192c	Dropped html_ prefix from HTML lexer specs	2015-05-23 09:48:45 +02:00
Yorick Peterse	f587b49406	Move HTML lexer specs into spec/oga/html/lexer	2015-05-23 09:47:49 +02:00
Yorick Peterse	04f431eee7	Add HTML lexer large input timing benchmark	2015-05-23 06:54:15 +02:00
Yorick Peterse	ce10ac779d	Move HTML lexer benchmarks to a separate directory	2015-05-23 06:53:11 +02:00
Yorick Peterse	3c6263d8de	Updated list of elements that close <p> tags	2015-05-21 21:11:41 +02:00

1 2 3 4 5 ...

1085 Commits All Branches Search

1085 Commits

All Branches